LLM API prices across 70+ providers: cheapest offer, comparisons, history and cost estimates.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
One-click editor setup isnβt available for this listing yet β we donβt have a confirmed install command, and weβd rather show nothing than point your editor at the wrong package or host. Follow the projectβs own setup instructions, linked above.
InferIndex finds the cheapest provider for a given LLM, across 70+ tracked pricing sources β direct provider APIs and aggregators/routers β with price history and promotion detection.
Live API: https://api.inferindex.dev
Service status: https://status.inferindex.dev (external uptime monitoring)
This repository holds the public documentation, examples and (eventually) a client SDK for that API. The backend itself β collection, parsers, data model β lives in a separate private repository.
Example response as of 2026-09-18 β prices, providers and statuses change.
Prices are in USD per million tokens, converted at the day's ECB rate when the source publishes another
currency. blended_per_1M is (3 Γ input + output) / 4, a rough proxy for a typical chat workload.
Each offer appears once, at its cheapest source: via is "direct" or the aggregator it comes from, and
also_via lists other sources selling the same offer. promo is false unless a promotion was detected, in
which case confidence and price_before_promo are filled in.
Early / actively developed. The API is public and free to use, without authentication, but there is no uptime guarantee, no versioned API contract yet, and endpoints or response shapes may change. If you build something on it, expect to adjust when they do. Feedback and issues are welcome.
Service status: https://status.inferindex.dev (external uptime monitoring). It shows current and past availability; it is not an uptime guarantee.
Provider pricing pages disagree, change without notice, and rarely show what resellers actually charge for the same model. InferIndex collects prices directly from 70+ sources on their own schedule (from hourly to daily depending on the source), stores every change, and answers "who is cheapest for model X right now" from that history β client requests never trigger a live call to a provider.
Cache-Control: public, max-age=300 (5 minutes) (except /health/ready, never cached). On the custom domain, responses are additionally held in Cloudflare's
edge cache for the same 5 minutes./cheapest, /history and /resellers, counting only
requests not answered from cache. An optional API key (x-api-key header) raises this to 300 requests per
minute per key; free keys will be offered, self-service sign-up isn't available yet. An invalid key returns
401. Over the limit, requests get 429. The MCP server has its own limits, see
Use with AI assistants.prompt_tokens, output_tokens, cached_ratio and requests_per_day to /cheapest
or /resellers to get an estimated cost per request and per month for each offer, and
sort=estimated_cost to rank by it. See docs/api.md.flex, batch β lower priority, variable latency, in
exchange for a lower price) are excluded from /cheapest and /resellers unless you ask for them with
include_tiers=flex,batch. The number hidden is reported in hidden_tiers.Data corrections are listed in the changelog.
InferIndex is a comparison tool, not a source of truth. Treat every price as "what our collector last saw", not as a guarantee of what you will be billed:
| Route | What it does |
|---|---|
GET /cheapest?model=deepseek/deepseek-v3.2 | Cheapest offer across all active sources, and every matching offer deduplicated and sorted |
GET /resellers?model=deepseek/deepseek-v3.2 | Current prices across every reseller (direct and via aggregators), one line per source, paginated beyond 100 offers |
GET /history?model=deepseek/deepseek-v3.2&days=30&granularity=day | Price history, paginated by cursor |
GET /models?search=deepseek | Search tracked models |
GET /trending | Up to 10 models currently getting attention, recomputed about every 6 hours |
GET /index?series=market | Weekly market index, every Monday from 2026-10-05 β see the methodology |
GET /health/live | Liveness only: the service is up, no database access |
GET /health/ready | Readiness: data freshness and scheduler health, see below |
/mcp | MCP server for AI assistants, see Use with AI assistants |
/cheapest filterssort=blended\|input\|output β sort order (default blended)min_context=100000 β minimum context windowquantization=fp8,bf16 β accepted quantizationsmin_uptime=95 β minimum 30-minute uptime (%), as observed by an aggregator (not measured by InferIndex)tools=true, json=true, vision=true β only offers declared to support tool calling, JSON output, image inputregion=eu β only offers processed in that region (eu, us, cn, uk, ch, sg, id, my, vn, kr, jp, in, ca, au), per the providers' official documentsno_training=true β only offers whose providers state they don't train on promptsno_waitlist=true β only providers whose signup is open to everyonestrict=true β also drop offers for which the filtered data is unknown (by default they're kept and flagged
in unverified, counted in filters_unknown)include_tiers=flex,batch β include degraded service tiers (excluded by default, see above)explain=true β list every excluded offer with its reason codes (an explanation summary is always included)Offers also carry conditions (data region, retention, training on prompts, and what it takes to open an
account β each with its official source link), quantization_source (whether the compute precision is declared
by the source or inferred), and, on /resellers, reliability from the providers' official status pages. See
docs/api.md.
/historygranularity=raw returns one row per price change; day / week (weeks start Monday) return one point per
period and per offer (provider, quantization, variant) with min, max and last price. Use days or from/to
for a period, or at=2026-09-15 for the prices in force at a given moment. Prices are converted to USD at the
ECB rate of the day of the price.
Model search (/models):
Example response as of 2026-09-15 β prices, providers and statuses change.
/health/readyExample response as of 2026-09-15 β prices, providers and statuses change.
status is ok, degraded (a non-critical part is having trouble β see the degraded array for which) or
unavailable (no usable price data). HTTP 200 for ok/degraded, 503 for unavailable. Useful as an
uptime-monitor target if you depend on this API.
Full route and parameter reference: docs/api.md.
InferIndex is also available as a remote MCP server, so an AI assistant can look up prices for you:
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/inferindex)<a href="https://allmcps.com/mcp/inferindex"><img src="https://allmcps.com/api/badge/inferindex?style=directory" alt="InferIndex on AllMCPs" /></a>