The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the TokCalc MCP Server listing page.
Plan your LLM deployment before you rent the GPUs.

Switch model → multi-GPU → long-context capacity planner → Build vs Buy → Reference catalog → Live GPU pricing
tokcalc turns your LLM traffic, context length, latency SLOs, cache behavior, and model choice into a defensible serving topology and cost plan — with transparent formulas and cited benchmarks.
It's the engineering-grade pre-deployment decision layer for LLM inference. Not another static "tokens per second" calculator.
tokcalc.vercel.app — no signup, no tracking, no paywall.
Pick a model, a GPU, and a workload. Get an instant capacity plan:
The market has dozens of "tokens per second" calculators and self-host-vs-API break-even tools (induwara.lk, gigagpu, kickllm, cloudparity, curlscape, profitable.ai). None of them are unified capacity planners — and none show their math.
"Can I serve Qwen 2.5 72B at 128K context on 2× H100 with 20 concurrent users?"
"How many H200s do I need for 1,000 req/min with P95 TTFT < 2s?"
"Does FP8 or AWQ save more money once quality, KV cache, and engine support are included?"
"At what daily volume does an H100 beat GPT-4o pricing?"
"What's the cheapest H100 right now across Azure / AWS / GCP / Vast.ai?"
"What happens to cost and latency if an agent makes 8 model calls, has 3 tool calls, and its context grows by 5K tokens each turn?"
"Would prefix caching, continuous batching, or PD disaggregation save more for this workload?"
| Feature | What it computes |
|---|---|
| Model fit / VRAM | Will the model + KV cache fit in the GPU's memory? |
| Throughput | Decode tok/s (per-stream) + aggregate (batched) + prefill tok/s |
| Latency split | Time-to-first-token (= prefill) + inter-token latency (= decode) |
| Continuous batching | User-tunable 1.0–4× multiplier (cited 1.5–4× SOSP range) |
| Reasoning tokens | Hidden reasoning budget added to billed output (o1 / R1 / Claude thinking) |
| Prompt caching | Self-hosted vLLM APC + Anthropic 5m/1h TTL + OpenAI 50%-off cached tokens |
| Speculative decoding | User-tunable 1.2–4× boost factor |
| Multi-GPU TP | 1× → 8× tensor parallel with NVLink efficiency factor |
| Long-context capacity | KV memory + max concurrency + prefill time at 4K → 1M context |
| Topology recommendation | Single GPU → TP×2 → TP×4 → TP×8 → TP×8 + Context Parallel (RingAttention) |
| Cost economics | GPU $/hr → $/M output tokens → $/request → monthly cost |
| Observed benchmark calibration | Paste vLLM/SGLang/TRT-LLM JSON → see formula accuracy verdict (validated / underestimated / overestimated) |
Independent calculator (separate state) that compares:
6 sub-tables — fully transparent, every record source-linked where available:
This is the formula the Perplexity research brief called "the most important tokcalc should visibly expose":
$$ \text{KV bytes/request} = 2 \cdot L \cdot T \cdot H_{\text{kv}} \cdot D_h \cdot B $$
For dense attention, prefill cost grows superlinearly with context length:
$$ \text{prefill FLOPs} = \underbrace{2 \cdot N \cdot T}{\text{linear}} + \underbrace{T^2 \cdot H{\text{kv}} \cdot D_h \cdot L}_{\text{attention}} $$
tokcalc shows you:
The 🔴 Live pricing sub-tab in Reference fetches real-time GPU prices from 4 providers in parallel via Promise.allSettled:
| Provider | API | Auth | Cache TTL |
|---|---|---|---|
| Azure | Retail Prices API | None (public) | 24h |
| AWS | EC2 bulk pricing file | None (public) | 24h |
| GCP | Cloud Billing Catalog API | GCP_API_KEY env var | 24h |
| Vast.ai | Marketplace bundles API | None (public) | 5m (spot prices change rapidly) |
Each provider shows a LIVE badge with timestamp + cache state. The comparison table shows the cheapest price per GPU across all 4 providers (highlighted in emerald) + per-provider breakdown.
If any provider fails (e.g., GCP_API_KEY not set), the others still work — graceful degradation per provider.
Every number above comes from a formula you can inspect. No black boxes.
$$ \text{decode tok/sec} \approx \frac{\text{HBM BW} \cdot \eta_{\text{mem}} \cdot \text{quant_eff}}{\text{model size}} $$
Where:
HBM BW = GPU memory bandwidth (e.g., 3350 GB/s for H100 SXM)η_mem = 0.65 = typical real-world memory utilization (35% overhead)quant_eff = dequantization efficiency multiplier (1.0 for FP16, 1.5 for FP8 on H100, 0.85 for INT4)model size = active_params × bytes_per_param (uses ACTIVE params for MoE, not total)Refs: PagedAttention paper (arxiv.org/abs/2309.06180)
$$ \text{prefill tok/sec} \approx \frac{\text{GPU FLOPS} \cdot \eta_{\text{compute}}}{2 \cdot \text{active params}} $$
Where:
GPU FLOPS = dense FP16/BF16 TFLOPS (sparse values not used)η_compute = 0.50 = typical compute utilizationFor long context (>32K), the superlinear attention correction above applies.
$$ \text{aggregate tok/sec} = \text{decode tok/sec} \cdot \text{batch size} \cdot \text{continuous batching multiplier} $$
Critical caveat: There is no universal continuous batching multiplier. vLLM reported 14–24× vs HF Transformers (extreme), 2.2–2.5× vs TGI. SOSP paper finds 2–4× typical vs FasterTransformer/Orca. tokcalc defaults to a conservative 1.5× and lets you tune.
Refs:
For shared prefix of length $T_p$, suffix of length $T_u$, output $O$, cache hit rate $h$:
$$ \text{API input cost} = N \cdot \left[ (1-h) \cdot T_p \cdot P_{\text{write}} + h \cdot T_p \cdot P_{\text{read}} + T_u \cdot P_{\text{input}} \right] $$
Anthropic multipliers (verified 2025-2026):
OpenAI: cached input discounted 50%, no separate write fee.
Refs:
$$ \text{total needed} = \text{model weights} + (\text{KV per request} \cdot \text{batch size}) $$
Walk the smallest topology that fits:
Refs: RingAttention paper
$$ \text{self-host $/M tokens} = \frac{\text{GPU $/hr}}{3600 \cdot \text{effective tok/s} \cdot \text{utilization}} \cdot 10^6 $$
$$ \text{break-even req/day} = \frac{\text{monthly self-host cost}}{30 \cdot \text{API cost per request}} $$
The decisive term is effective utilization — not peak throughput. A GPU running at 10% utilization pays 10× more per token than the theoretical minimum.
| Capability | induwara / techfuelhq / pcmasterstudio | gigagpu / kickllm / cloudparity | HF Open LLM Leaderboard / MLPerf | tokcalc |
|---|---|---|---|---|
| Model × GPU × quant tok/s | ✓ | some | some | ✓ |
| Model-fit / VRAM | basic | rare | rare | ✓ |
| KV-cache by context + concurrency | — | — | implicit | ✓ |
| Prefill vs decode split (TTFT/ITL) | — | — | engine-specific | ✓ |
| Continuous batching / paged attention | — | — | docs only | ✓ |
| Long-context (128K–1M) planning | — | — | — | ✓ |
| Topology recommendation (TP/CP) | — | — | partial | ✓ |
| Prompt-cache economics | — | partial API only | — | ✓ |
| Reasoning tokens (o1/R1/Claude thinking) | — | — | — | ✓ |
| API vs self-host break-even | some | ✓ | — | ✓ |
| Live cloud GPU pricing (4 providers) | — | — | — | ✓ |
| Observed benchmark calibration | — | — | — | ✓ |
| MCP server for AI agents | — | — | — | ✓ |
| Public hosted MCP endpoint + self-serve keys | — | — | — | ✓ |
| Transparent formulas / open source | mixed | usually no | mixed | ✓ |
| Cited benchmark evidence per config | rare | rare | ✓ (not planning) | ✓ (in progress) |
| Shareable URL per config | — | — | — | ✓ |
| Copy result as Markdown | — | — | — | ✓ |
tokcalc.vercel.app/api/mcp with bearer auth + Upstash Redis rate limitingtokcalc.vercel.app/mcp (email → instant key)/compare/h100-vs-h200, /compare/gguf-q4-k-m-vs-q5-k-m, /self-host-vs-openai-api, /mcp (install docs)tokcalc/plan PR comment)inputSchema in tools/list response (zod-to-json-schema serialization issue)tokcalc is open core — the calculator and catalog are open source; the cloud/data/team features are paid.
| Asset | License | Notes |
|---|---|---|
| Source code | Apache 2.0 | This repo. Free to use, modify, distribute |
| Model/GPU/quant catalog | CC0 1.0 | Public domain data. Anyone can use, no attribution required |
| Benchmark provenance data | CC-BY-SA 4.0 | Anyone can use, but must attribute + share-alike |
| Documentation | CC-BY 4.0 | Attribution required if copied |
| "tokcalc" name + logo | Trademark | Even without formal registration, common-law rights apply |
| Cloud SaaS layer | Proprietary | Real-time pricing API, benchmark DB, team workspaces (coming soon) |
We welcome contributions! See CONTRIBUTING.md for:
tokcalc ships an MCP (Model Context Protocol) server that lets AI agents (Cursor, Claude Desktop, Cline) call tokcalc during design reviews.
npm package: @tokcalc/mcp-server (latest: v0.2.0)
Public endpoint: https://tokcalc.vercel.app/api/mcp (Streamable HTTP, bearer auth, rate-limited)
Install docs: https://tokcalc.vercel.app/mcp
| Tool | What it does |
|---|---|
estimate_capacity | VRAM/KV/throughput/latency/cost for one config |
compare_gpus | Ranked GPU comparison for one workload |
recommend_topology | TP/CP topology recommendation |
estimate_api_vs_self_host | Break-even analysis |
list_models | Discover supported model IDs |
list_gpus | Discover supported GPU IDs |
get_mlperf_benchmarks | Curated MLPerf Inference v4.1 audited reference configs |
All tools are read-only — no side effects, no cloud credentials, no deployments.
Point Cursor or Claude Desktop at the public endpoint. Get an instant API key at /mcp:
Features:
crypto.timingSafeEqual)/mcp, get instant key (5 per IP per day limit)No API key needed for stdio (local install). Cursor spawns the process via npx.
Set MCP_API_KEY, UPSTASH_REDIS_REST_URL, UPSTASH_REDIS_REST_TOKEN env vars for auth + rate limiting.
"I need to serve Llama 3.3 70B at 32K context for 50 concurrent users. What GPU topology do you recommend, and how much will it cost per month?"
The agent calls list_models → list_gpus → recommend_topology → estimate_capacity → returns a structured plan with VRAM, throughput, latency, cost, and confidence.
@modelcontextprotocol/sdk v1.30.1 (Streamable HTTP transport, stateless mode)crypto.timingSafeEqual (constant-time bearer key validation)@upstash/ratelimit + @upstash/redis (sliding window, 30/min per IP, 120/min per key)tokcalc builds on the work of:
Every formula has a citation. Every model/GPU/quant entry has a source URL where available. If you spot an unsourced claim, please open an issue.
Live demo · MCP install docs · Public MCP endpoint · npm package · Contributing · License · Code of Conduct
Made with care by the tokcalc community. Apache 2.0 licensed.