The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Kosha Model Discovery listing page.
Tells your agent — or your code — which model to use and what it costs.
kosha discovers models across 45 providers and local runtimes, finds your API keys wherever they already live (env vars, Claude CLI, Codex, gcloud ADC, AWS SSO), fills in pricing and context limits, and answers questions like the cheapest model with tool use and 128k context that I actually hold a key for. It ships as a TypeScript library, a CLI, an HTTP API, an OpenAI-compatible proxy with a spend ledger, and an MCP server.
It works with no API keys at all — discovery falls back to the public models.dev and
LiteLLM catalogs, so kosha list is useful on a fresh machine.
Requires Node.js 22+.
Every command takes --json. kosha --help lists the rest.
After each discovery, a stable v1 manifest lands at ~/.kosha/registry.json:
A weekly discovery run publishes a full snapshot — every provider, model, price and limit kosha can see without your keys — at a stable URL:
It is a release asset rather than a file in the repository: at ~2.8 MB growing
with every provider added, committing it weekly would put roughly 150 MB of
already-stale data a year into a repo people are meant to clone. Dated
snapshot-YYYY-MM-DD pre-releases keep a short trail for diffing, pruned to the
two most recent, and an older one is only removed once a newer one exists.
Parameters and response shapes: docs/api.md.
kosha serve also exposes an OpenAI-compatible endpoint at /proxy/v1. Point any OpenAI SDK at it; the proxy resolves the model or alias, picks a provider you hold credentials for, injects the upstream key, forwards the request, and writes a row to the spend ledger.
kosha:cheapest filter syntax (comma-separated, combinable):
| Filter | Example | Meaning |
|---|---|---|
| capability | tool_use, vision | model must have this tag |
<N>k | 128k, 200k | minimum context window |
provider:<id> | provider:groq | pin to a specific provider |
kosha:fastest, kosha:reliable, and kosha:balanced take the same filters and rank on observed latency and circuit-breaker state instead of price.
Every response carries x-kosha-model, x-kosha-provider, x-kosha-requested, x-kosha-attempt-chain, and x-kosha-estimated-cost-usd; non-streaming responses add x-kosha-actual-cost-usd when the upstream returned a usage block.
What the proxy can forward:
| Upstream wire format | Providers | Support |
|---|---|---|
| OpenAI-compatible | OpenAI, Ollama, OpenRouter, Vercel, Groq, Together, Fireworks, DeepInfra, … | passthrough, streaming included |
| Anthropic Messages | Anthropic | translated: streaming, tools, image_url, response_format, reasoning_effort; audio input and non-function tools fail over to an OpenAI-compatible route for the same model |
| Cloud SDKs | Google, Bedrock, Vertex | discovery only, not proxied yet |
| Non-chat / other wire | TypeSafe, Thinking Machines | discovery only — TypeSafe's System One endpoint is not a chat API, and Tinker's Anthropic-wire path differs from Anthropic's own |
Defaults that matter before you expose it: the server binds 127.0.0.1. Pass --host 0.0.0.0 (or KOSHA_HOST) to listen on a network interface, and set KOSHA_PROXY_TOKEN so /proxy/* and POST /api/refresh require Authorization: Bearer <token> or x-kosha-token. KOSHA_MONTHLY_BUDGET_USD caps spend per calendar month. Reference: docs/api.md, docs/operations.md.
kosha-mcp serves the registry over the Model Context Protocol on stdio, so an agent can call kosha_query_models, kosha_cheapest_model, kosha_ranked_routes, kosha_model_detail, kosha_model_routes, kosha_resolve_alias, kosha_provider_health, and kosha_context_strategy without an HTTP server.
It is published to the MCP registry
as io.github.sriinnu/kosha-discovery, so clients that read the registry can
install it without a manual command. The manifest is server.json.
Every provider key is optional. With no credentials at all the server still answers from the public models.dev and LiteLLM catalogs plus a curated offline list, so it is useful on a fresh machine.
Tools and protocol details: docs/mcp.md.
45 providers. Each has a descriptor in src/provider-catalog.ts; most OpenAI-compatible
ones are driven from GENERIC_OPENAI_PROVIDERS in src/discovery/generic-openai.ts
rather than a hand-written class.
| Provider | Discovery | Credential sources |
|---|---|---|
| Anthropic | GET /v1/models (context, output cap, capabilities read from the API) | ANTHROPIC_API_KEY, Claude CLI, Codex CLI |
| OpenAI | GET /v1/models | OPENAI_API_KEY, GitHub Copilot tokens |
GET /v1beta/models | GOOGLE_API_KEY, GEMINI_API_KEY, Gemini CLI, gcloud | |
| AWS Bedrock | SDK → CLI → static list | AWS_ACCESS_KEY_ID, ~/.aws/credentials, SSO, IAM |
| Vertex AI | API + gcloud | GOOGLE_APPLICATION_CREDENTIALS, ADC |
| Ollama, llama.cpp, LM Studio, vLLM | local HTTP API | none |
| OpenRouter | API | OPENROUTER_API_KEY (optional; unauthenticated is rate-limited) |
| Vercel AI Gateway | GET /v1/models | AI_GATEWAY_API_KEY, VERCEL_OIDC_TOKEN (discovery works without; execution needs one) |
| NVIDIA, Together, Fireworks, Groq, Cerebras, Cohere, DeepInfra, Perplexity | OpenAI-compatible API | <PROVIDER>_API_KEY |
| DeepSeek, Mistral, Moonshot (Kimi), GLM (Zhipu), Z.AI, MiniMax | OpenAI-compatible API | <PROVIDER>_API_KEY |
| xAI (Grok) | GET /v1/models; Grok Imagine split into image / video | XAI_API_KEY |
| TypeSafe (System One / Jev) | GET /v1/models — returns judgment models, not chat | TYPESAFE_API_KEY, JEV_API_KEY |
| Thinking Machines (Inkling) | Anthropic-wire endpoint, no model list — public catalog only | TINKER_API_KEY |
| Alibaba Model Studio (Qwen), Volcengine Ark (Doubao), Inception (Mercury), AI21 (Jamba), Upstage (Solar), StepFun | OpenAI-compatible API | DASHSCOPE_API_KEY, ARK_API_KEY, INCEPTION_API_KEY, AI21_API_KEY, UPSTAGE_API_KEY, STEPFUN_API_KEY |
| Baseten, Nebius Token Factory, Novita AI, SiliconFlow, Hugging Face, Ollama Cloud | OpenAI-compatible API | <PROVIDER>_API_KEY, HF_TOKEN |
Several providers run separate hosts for international and mainland-China traffic, with separate keys and separate price sheets — Qwen 2.5 72B is $1.40/M input internationally against $0.574/M in China. Merging them would make a model's price depend on which host answered last, so each region is its own provider:
| International | China | Differs in |
|---|---|---|
moonshot (api.moonshot.ai) | moonshot-cn (api.moonshot.cn) | host, key |
minimax (api.minimax.io) | minimax-cn (api.minimaxi.com) | host, key |
alibaba (dashscope-intl) | alibaba-cn (dashscope) | host, key, pricing |
siliconflow (.com) | siliconflow-cn (.cn) | host, key, pricing |
stepfun (api.stepfun.ai) | stepfun-cn (api.stepfun.com) | host, key, pricing |
zai (api.z.ai) | glm (open.bigmodel.cn) | host, key, pricing |
kosha routes <model> lists every region a model is served from, so you can
compare prices across them directly.
Without a key, providers fall back to the public models.dev + LiteLLM catalog, then to a curated static list, so kosha list works on a fresh machine. Exact env var names: docs/credentials.md.
Most models are chat, but mode also covers embedding, image, video,
audio, moderation, rerank, and judgment. A judgment model — TypeSafe's
System One family — answers a question with a typed value (a choice and its
probability distribution, a probability, or a score on described levels) instead
of generating text. It carries no chat capability on purpose: routing a prompt
to one would be a category error.
Promise.allSettled); each returns normalized ModelCards. A failing provider is recorded in discoveryErrors() and doesn't block the others.pricingSource on each card says which.~/.kosha/cache/ (24 h TTL) and exported as a versioned snapshot at ~/.kosha/registry.json for other tools to read.kosha:<strategy>[filters] selector against the registry, ranks candidate routes, forwards with failover, and records estimated and reconciled cost in ~/.kosha/ledger-YYYY-MM.jsonl.Details: docs/architecture.md, docs/resilience.md.
For an OpenAI-compatible provider — most of them — it is two table entries:
PROVIDER_CATALOG in src/provider-catalog.ts (id, base URL, transport, credential env vars).GENERIC_OPENAI_PROVIDERS in src/discovery/generic-openai.ts. It registers its own discoverer, resolves credentials from the descriptor's credentialEnvVars, and falls back to the public catalog when there is no key.src/discovery/modelsdev-seed.ts and litellm-seed.ts so keyless discovery works. Check the slug — GLM is published as zhipuai, and a missing mapping means no keyless models at all.test/discovery/generic-openai.test.ts and document the env vars in docs/credentials.md.Write a discoverer class only when classification genuinely needs code — per-route pricing (OpenRouter), a non-OpenAI response envelope (TypeSafe), or origin remapping (Vercel):
src/discovery/<provider>.ts extending BaseDiscoverer or OpenAICompatibleDiscoverer; fetchJSON and makeCard are provided.src/discovery/index.ts and add a factory entry to DISCOVERER_REGISTRY.src/credentials/resolver.ts only for a bespoke search order (CLI files, OAuth, ADC, SSO). A plain API key in env vars needs no branch — the descriptor covers it.test/discovery/<provider>.test.ts (mock fetch; see typesafe.test.ts).| Credentials | Env vars, CLI tools, and config files for every provider |
| CLI | Commands, flags, examples |
| HTTP API | Endpoints, parameters, response schemas |
| MCP server | Tools, protocol negotiation, client setup |
| Configuration | Aliases, routing, enrichment, programmatic config |
| Architecture | Discovery flow, module map, adding providers |
| Resilience | Circuit breakers, stale cache, health |
| Operations | Deployment sizing, metrics, spend ledger, recovery recipes |
| Security | Threat catalogue, runtime scanning, pre-commit hook |
| Discovery Plane v1 | Stable daemon contract (deltas, SSE watch, binding hints) |
version in package.json and date the [Unreleased] section in CHANGELOG.md; merge that as a PR.The workflow checks that the tag matches package.json, runs lint / build / test, publishes to npm with provenance, and creates the GitHub Release. Publishing authenticates through npm trusted publishing (OIDC) when a trusted publisher is configured for this repo and workflow on npmjs.com, or through an NPM_TOKEN repository secret.
MIT