LLM capacity planning tools to estimate VRAM, compute, latency, and GPU topology for agents.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
One-click editor setup isnβt available for this listing yet β we donβt have a confirmed install command, and weβd rather show nothing than point your editor at the wrong package or host. Follow the projectβs own setup instructions, linked above.
Plan your LLM deployment before you rent the GPUs.

Switch model β multi-GPU β long-context capacity planner β Build vs Buy β Reference catalog β Live GPU pricing
tokcalc turns your LLM traffic, context length, latency SLOs, cache behavior, and model choice into a defensible serving topology and cost plan β with transparent formulas and cited benchmarks.
It's the engineering-grade pre-deployment decision layer for LLM inference. Not another static "tokens per second" calculator.
tokcalc.vercel.app β no signup, no tracking, no paywall.
Pick a model, a GPU, and a workload. Get an instant capacity plan:
The market has dozens of "tokens per second" calculators and self-host-vs-API break-even tools (induwara.lk, gigagpu, kickllm, cloudparity, curlscape, profitable.ai). None of them are unified capacity planners β and none show their math.
"Can I serve Qwen 2.5 72B at 128K context on 2Γ H100 with 20 concurrent users?"
"How many H200s do I need for 1,000 req/min with P95 TTFT < 2s?"
"Does FP8 or AWQ save more money once quality, KV cache, and engine support are included?"
"At what daily volume does an H100 beat GPT-4o pricing?"
"What's the cheapest H100 right now across Azure / AWS / GCP / Vast.ai?"
"What happens to cost and latency if an agent makes 8 model calls, has 3 tool calls, and its context grows by 5K tokens each turn?"
"Would prefix caching, continuous batching, or PD disaggregation save more for this workload?"
| Feature | What it computes |
|---|---|
| Model fit / VRAM | Will the model + KV cache fit in the GPU's memory? |
| Throughput | Decode tok/s (per-stream) + aggregate (batched) + prefill tok/s |
| Latency split | Time-to-first-token (= prefill) + inter-token latency (= decode) |
| Continuous batching | User-tunable 1.0β4Γ multiplier (cited 1.5β4Γ SOSP range) |
| Reasoning tokens | Hidden reasoning budget added to billed output (o1 / R1 / Claude thinking) |
| Prompt caching | Self-hosted vLLM APC + Anthropic 5m/1h TTL + OpenAI 50%-off cached tokens |
| Speculative decoding | User-tunable 1.2β4Γ boost factor |
| Multi-GPU TP | 1Γ β 8Γ tensor parallel with NVLink efficiency factor |
| Long-context capacity | KV memory + max concurrency + prefill time at 4K β 1M context |
| Topology recommendation | Single GPU β TPΓ2 β TPΓ4 β TPΓ8 β TPΓ8 + Context Parallel (RingAttention) |
| Cost economics | GPU $/hr β $/M output tokens β $/request β monthly cost |
| Observed benchmark calibration | Paste vLLM/SGLang/TRT-LLM JSON β see formula accuracy verdict (validated / underestimated / overestimated) |
Independent calculator (separate state) that compares:
6 sub-tables β fully transparent, every record source-linked where available:
This is the formula the Perplexity research brief called "the most important tokcalc should visibly expose":
$$ \text{KV bytes/request} = 2 \cdot L \cdot T \cdot H_{\text{kv}} \cdot D_h \cdot B $$
For dense attention, prefill cost grows superlinearly with context length:
$$ \text{prefill FLOPs} = \underbrace{2 \cdot N \cdot T}{\text{linear}} + \underbrace{T^2 \cdot H{\text{kv}} \cdot D_h \cdot L}_{\text{attention}} $$
tokcalc shows you:
The π΄ Live pricing sub-tab in Reference fetches real-time GPU prices from 4 providers in parallel via Promise.allSettled:
| Provider | API | Auth | Cache TTL |
|---|---|---|---|
| Azure | Retail Prices API | None (public) | 24h |
| AWS | EC2 bulk pricing file | None (public) | 24h |
| GCP | Cloud Billing Catalog API | GCP_API_KEY env var | 24h |
| Vast.ai | Marketplace bundles API | None (public) | 5m (spot prices change rapidly) |
Each provider shows a LIVE badge with timestamp + cache state. The comparison table shows the cheapest price per GPU across all 4 providers (highlighted in emerald) + per-provider breakdown.
If any provider fails (e.g., GCP_API_KEY not set), the others still work β graceful degradation per provider.
Every number above comes from a formula you can inspect. No black boxes.
$$ \text{decode tok/sec} \approx \frac{\text{HBM BW} \cdot \eta_{\text{mem}} \cdot \text{quant_eff}}{\text{model size}} $$
Where:
HBM BW = GPU memory bandwidth (e.g., 3350 GB/s for H100 SXM)Ξ·_mem = 0.65 = typical real-world memory utilization (35% overhead)quant_eff = dequantization efficiency multiplier (1.0 for FP16, 1.5 for FP8 on H100, 0.85 for INT4)model size = active_params Γ bytes_per_param (uses ACTIVE params for MoE, not total)Refs: PagedAttention paper (arxiv.org/abs/2309.06180)
$$ \text{prefill tok/sec} \approx \frac{\text{GPU FLOPS} \cdot \eta_{\text{compute}}}{2 \cdot \text{active params}} $$
Where:
GPU FLOPS = dense FP16/BF16 TFLOPS (sparse values not used)Ξ·_compute = 0.50 = typical compute utilizationFor long context (>32K), the superlinear attention correction above applies.
$$ \text{aggregate tok/sec} = \text{decode tok/sec} \cdot \text{batch size} \cdot \text{continuous batching multiplier} $$
Critical caveat: There is no universal continuous batching multiplier. vLLM reported 14β24Γ vs HF Transformers (extreme), 2.2β2.5Γ vs TGI. SOSP paper finds 2β4Γ typical vs FasterTransformer/Orca. tokcalc defaults to a conservative 1.5Γ and lets you tune.
Refs:
For shared prefix of length $T_p$, suffix of length $T_u$, output $O$, cache hit rate $h$:
$$ \text{API input cost} = N \cdot \left[ (1-h) \cdot T_p \cdot P_{\text{write}} + h \cdot T_p \cdot P_{\text{read}} + T_u \cdot P_{\text{input}} \right] $$
Anthropic multipliers (verified 2025-2026):
OpenAI: cached input discounted 50%, no separate write fee.
Refs:
$$ \text{total needed} = \text{model weights} + (\text{KV per request} \cdot \text{batch size}) $$
Walk the smallest topology that fits:
Refs: RingAttention paper
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/tokcalc-mcp-server)<a href="https://allmcps.com/mcp/tokcalc-mcp-server"><img src="https://allmcps.com/api/badge/tokcalc-mcp-server?style=directory" alt="TokCalc MCP Server on AllMCPs" /></a>