The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the InferIndex listing page.
InferIndex finds the cheapest provider for a given LLM, across 70+ tracked pricing sources — direct provider APIs and aggregators/routers — with price history and promotion detection.
Live API: https://api.inferindex.dev
Service status: https://status.inferindex.dev (external uptime monitoring)
This repository holds the public documentation, examples and (eventually) a client SDK for that API. The backend itself — collection, parsers, data model — lives in a separate private repository.
Example response as of 2026-09-18 — prices, providers and statuses change.
Prices are in USD per million tokens, converted at the day's ECB rate when the source publishes another
currency. blended_per_1M is (3 × input + output) / 4, a rough proxy for a typical chat workload.
Each offer appears once, at its cheapest source: via is "direct" or the aggregator it comes from, and
also_via lists other sources selling the same offer. promo is false unless a promotion was detected, in
which case confidence and price_before_promo are filled in.
Early / actively developed. The API is public and free to use, without authentication, but there is no uptime guarantee, no versioned API contract yet, and endpoints or response shapes may change. If you build something on it, expect to adjust when they do. Feedback and issues are welcome.
Service status: https://status.inferindex.dev (external uptime monitoring). It shows current and past availability; it is not an uptime guarantee.
Provider pricing pages disagree, change without notice, and rarely show what resellers actually charge for the same model. InferIndex collects prices directly from 70+ sources on their own schedule (from hourly to daily depending on the source), stores every change, and answers "who is cheapest for model X right now" from that history — client requests never trigger a live call to a provider.
Cache-Control: public, max-age=300 (5 minutes) (except /health/ready, never cached). On the custom domain, responses are additionally held in Cloudflare's
edge cache for the same 5 minutes./cheapest, /history and /resellers, counting only
requests not answered from cache. An optional API key (x-api-key header) raises this to 300 requests per
minute per key; free keys will be offered, self-service sign-up isn't available yet. An invalid key returns
401. Over the limit, requests get 429. The MCP server has its own limits, see
Use with AI assistants.prompt_tokens, output_tokens, cached_ratio and requests_per_day to /cheapest
or /resellers to get an estimated cost per request and per month for each offer, and
sort=estimated_cost to rank by it. See docs/api.md.flex, batch — lower priority, variable latency, in
exchange for a lower price) are excluded from /cheapest and /resellers unless you ask for them with
include_tiers=flex,batch. The number hidden is reported in hidden_tiers.Data corrections are listed in the changelog.
InferIndex is a comparison tool, not a source of truth. Treat every price as "what our collector last saw", not as a guarantee of what you will be billed:
| Route | What it does |
|---|---|
GET /cheapest?model=deepseek/deepseek-v3.2 | Cheapest offer across all active sources, and every matching offer deduplicated and sorted |
GET /resellers?model=deepseek/deepseek-v3.2 | Current prices across every reseller (direct and via aggregators), one line per source, paginated beyond 100 offers |
GET /history?model=deepseek/deepseek-v3.2&days=30&granularity=day | Price history, paginated by cursor |
GET /models?search=deepseek | Search tracked models |
GET /trending | Up to 10 models currently getting attention, recomputed about every 6 hours |
GET /index?series=market | Weekly market index, every Monday from 2026-10-05 — see the methodology |
GET /health/live | Liveness only: the service is up, no database access |
GET /health/ready | Readiness: data freshness and scheduler health, see below |
/mcp | MCP server for AI assistants, see Use with AI assistants |
/cheapest filterssort=blended\|input\|output — sort order (default blended)min_context=100000 — minimum context windowquantization=fp8,bf16 — accepted quantizationsmin_uptime=95 — minimum 30-minute uptime (%), as observed by an aggregator (not measured by InferIndex)tools=true, json=true, vision=true — only offers declared to support tool calling, JSON output, image inputregion=eu — only offers processed in that region (eu, us, cn, uk, ch, sg, id, my, vn, kr, jp, in, ca, au), per the providers' official documentsno_training=true — only offers whose providers state they don't train on promptsno_waitlist=true — only providers whose signup is open to everyonestrict=true — also drop offers for which the filtered data is unknown (by default they're kept and flagged
in unverified, counted in filters_unknown)include_tiers=flex,batch — include degraded service tiers (excluded by default, see above)explain=true — list every excluded offer with its reason codes (an explanation summary is always included)Offers also carry conditions (data region, retention, training on prompts, and what it takes to open an
account — each with its official source link), quantization_source (whether the compute precision is declared
by the source or inferred), and, on /resellers, reliability from the providers' official status pages. See
docs/api.md.
/historygranularity=raw returns one row per price change; day / week (weeks start Monday) return one point per
period and per offer (provider, quantization, variant) with min, max and last price. Use days or from/to
for a period, or at=2026-09-15 for the prices in force at a given moment. Prices are converted to USD at the
ECB rate of the day of the price.
Model search (/models):
Example response as of 2026-09-15 — prices, providers and statuses change.
/health/readyExample response as of 2026-09-15 — prices, providers and statuses change.
status is ok, degraded (a non-critical part is having trouble — see the degraded array for which) or
unavailable (no usable price data). HTTP 200 for ok/degraded, 503 for unavailable. Useful as an
uptime-monitor target if you depend on this API.
Full route and parameter reference: docs/api.md.
InferIndex is also available as a remote MCP server, so an AI assistant can look up prices for you:
https://mcp.inferindex.dev/mcp (Streamable HTTP)
(https://api.inferindex.dev/mcp also works, as an alias)io.github.InferIndex/inferindex.x-api-key header./mcp, at most 10 JSON-RPC messages per batch and 64 KB per
request (otherwise 400 or 413 with JSON-RPC error -32600). Over the rate limit: 429 with JSON-RPC error
-32000. Each tool call also counts toward the limit of the API route it uses (/cheapest, /resellers,
/history).| Tool | What it does |
|---|---|
search_models(query) | Find the exact id of a model |
cheapest(model, …) | Cheapest offers, with the same filters as /cheapest (min_context, tools, json, vision, region, no_training, no_waitlist, strict, include_tiers), optional usage (prompt_tokens, output_tokens, cached_ratio, requests_per_day) and limit |
compare_providers(model, sort, limit, …) | Every offer, one line per provider, with usage conditions and reliability |
price_history(model, days | from + to | at, granularity, provider, limit) | Offer price history, plus the lab's official prices |
estimate_cost(model, prompt_tokens, output_tokens, cached_ratio, requests_per_day) | Estimated cost per request and per month, sorted |
Every result includes api_url, the equivalent API call, so you can check or reuse it.
For Cline and other agents that install servers themselves, see llms-install.md.
Claude Code
With an API key, add --header "x-api-key: YOUR_KEY".
Claude Desktop / claude.ai: Settings → Connectors → Add custom connector, URL https://mcp.inferindex.dev/mcp.
Older Claude Desktop versions, in claude_desktop_config.json:
Cursor: one click, or add it by hand.
In ~/.cursor/mcp.json:
With an API key, add "headers": { "x-api-key": "YOUR_KEY" } next to "url".
Codex CLI
Or in ~/.codex/config.toml:
With an API key, add env_http_headers = { "x-api-key" = "INFERINDEX_API_KEY" } to that table and set the
INFERINDEX_API_KEY environment variable.
Documentation fixes and example contributions are welcome — open a pull request or an issue. This repository covers the public API surface only; the backend itself is not open source.
docs/): CC BY 4.0. Reuse and adapt it freely, with credit
to InferIndex.These licenses cover this repository only, not the InferIndex backend or its price data.