Tells your agent which model to use and what it costs, across 45 providers and local runtimes.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
One-click editor setup isnβt available for this listing yet β we donβt have a confirmed install command, and weβd rather show nothing than point your editor at the wrong package or host. Follow the projectβs own setup instructions, linked above.
Tells your agent β or your code β which model to use and what it costs.
kosha discovers models across 45 providers and local runtimes, finds your API keys wherever they already live (env vars, Claude CLI, Codex, gcloud ADC, AWS SSO), fills in pricing and context limits, and answers questions like the cheapest model with tool use and 128k context that I actually hold a key for. It ships as a TypeScript library, a CLI, an HTTP API, an OpenAI-compatible proxy with a spend ledger, and an MCP server.
It works with no API keys at all β discovery falls back to the public models.dev and
LiteLLM catalogs, so kosha list is useful on a fresh machine.
Requires Node.js 22+.
Every command takes --json. kosha --help lists the rest.
After each discovery, a stable v1 manifest lands at ~/.kosha/registry.json:
A weekly discovery run publishes a full snapshot β every provider, model, price and limit kosha can see without your keys β at a stable URL:
It is a release asset rather than a file in the repository: at ~2.8 MB growing
with every provider added, committing it weekly would put roughly 150 MB of
already-stale data a year into a repo people are meant to clone. Dated
snapshot-YYYY-MM-DD pre-releases keep a short trail for diffing, pruned to the
two most recent, and an older one is only removed once a newer one exists.
Parameters and response shapes: docs/api.md.
kosha serve also exposes an OpenAI-compatible endpoint at /proxy/v1. Point any OpenAI SDK at it; the proxy resolves the model or alias, picks a provider you hold credentials for, injects the upstream key, forwards the request, and writes a row to the spend ledger.
kosha:cheapest filter syntax (comma-separated, combinable):
| Filter | Example | Meaning |
|---|---|---|
| capability | tool_use, vision | model must have this tag |
<N>k | 128k, 200k | minimum context window |
provider:<id> | provider:groq | pin to a specific provider |
kosha:fastest, kosha:reliable, and kosha:balanced take the same filters and rank on observed latency and circuit-breaker state instead of price.
Every response carries x-kosha-model, x-kosha-provider, x-kosha-requested, x-kosha-attempt-chain, and x-kosha-estimated-cost-usd; non-streaming responses add x-kosha-actual-cost-usd when the upstream returned a usage block.
What the proxy can forward:
| Upstream wire format | Providers | Support |
|---|---|---|
| OpenAI-compatible | OpenAI, Ollama, OpenRouter, Vercel, Groq, Together, Fireworks, DeepInfra, β¦ | passthrough, streaming included |
| Anthropic Messages | Anthropic | translated: streaming, tools, image_url, response_format, reasoning_effort; audio input and non-function tools fail over to an OpenAI-compatible route for the same model |
| Cloud SDKs | Google, Bedrock, Vertex | discovery only, not proxied yet |
| Non-chat / other wire | TypeSafe, Thinking Machines | discovery only β TypeSafe's System One endpoint is not a chat API, and Tinker's Anthropic-wire path differs from Anthropic's own |
Defaults that matter before you expose it: the server binds 127.0.0.1. Pass --host 0.0.0.0 (or KOSHA_HOST) to listen on a network interface, and set KOSHA_PROXY_TOKEN so /proxy/* and POST /api/refresh require Authorization: Bearer <token> or x-kosha-token. KOSHA_MONTHLY_BUDGET_USD caps spend per calendar month. Reference: docs/api.md, docs/operations.md.
kosha-mcp serves the registry over the Model Context Protocol on stdio, so an agent can call kosha_query_models, kosha_cheapest_model, kosha_ranked_routes, kosha_model_detail, kosha_model_routes, kosha_resolve_alias, kosha_provider_health, and kosha_context_strategy without an HTTP server.
It is published to the MCP registry
as io.github.sriinnu/kosha-discovery, so clients that read the registry can
install it without a manual command. The manifest is server.json.
Every provider key is optional. With no credentials at all the server still answers from the public models.dev and LiteLLM catalogs plus a curated offline list, so it is useful on a fresh machine.
Tools and protocol details: docs/mcp.md.
45 providers. Each has a descriptor in src/provider-catalog.ts; most OpenAI-compatible
ones are driven from GENERIC_OPENAI_PROVIDERS in src/discovery/generic-openai.ts
rather than a hand-written class.
| Provider | Discovery | Credential sources |
|---|---|---|
| Anthropic | GET /v1/models (context, output cap, capabilities read from the API) | ANTHROPIC_API_KEY, Claude CLI, Codex CLI |
| OpenAI | GET /v1/models | OPENAI_API_KEY, GitHub Copilot tokens |
GET /v1beta/models | GOOGLE_API_KEY, GEMINI_API_KEY, Gemini CLI, gcloud | |
| AWS Bedrock | SDK β CLI β static list | AWS_ACCESS_KEY_ID, ~/.aws/credentials, SSO, IAM |
| Vertex AI | API + gcloud | GOOGLE_APPLICATION_CREDENTIALS, ADC |
| Ollama, llama.cpp, LM Studio, vLLM | local HTTP API | none |
| OpenRouter | API | OPENROUTER_API_KEY (optional; unauthenticated is rate-limited) |
| Vercel AI Gateway | GET /v1/models | AI_GATEWAY_API_KEY, VERCEL_OIDC_TOKEN (discovery works without; execution needs one) |
| NVIDIA, Together, Fireworks, Groq, Cerebras, Cohere, DeepInfra, Perplexity | OpenAI-compatible API | <PROVIDER>_API_KEY |
| DeepSeek, Mistral, Moonshot (Kimi), GLM (Zhipu), Z.AI, MiniMax | OpenAI-compatible API | <PROVIDER>_API_KEY |
| xAI (Grok) | GET /v1/models; Grok Imagine split into image / video | XAI_API_KEY |
| TypeSafe (System One / Jev) | GET /v1/models β returns judgment models, not chat | TYPESAFE_API_KEY, JEV_API_KEY |
| Thinking Machines (Inkling) | Anthropic-wire endpoint, no model list β public catalog only | TINKER_API_KEY |
| Alibaba Model Studio (Qwen), Volcengine Ark (Doubao), Inception (Mercury), AI21 (Jamba), Upstage (Solar), StepFun | OpenAI-compatible API | DASHSCOPE_API_KEY, ARK_API_KEY, INCEPTION_API_KEY, AI21_API_KEY, UPSTAGE_API_KEY, STEPFUN_API_KEY |
| Baseten, Nebius Token Factory, Novita AI, SiliconFlow, Hugging Face, Ollama Cloud | OpenAI-compatible API | <PROVIDER>_API_KEY, HF_TOKEN |
Several providers run separate hosts for international and mainland-China traffic, with separate keys and separate price sheets β Qwen 2.5 72B is $1.40/M input internationally against $0.574/M in China. Merging them would make a model's price depend on which host answered last, so each region is its own provider:
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/kosha-model-discovery)<a href="https://allmcps.com/mcp/kosha-model-discovery"><img src="https://allmcps.com/api/badge/kosha-model-discovery?style=directory" alt="Kosha Model Discovery on AllMCPs" /></a>