Skip to main content
AllMCPs
BrowseBestCategoriesStackCompareToolsGuidesBlog
Log in Submit MCP

Stay in the loop

Get new MCP servers and top picks in your inbox.

AllMCPs

The open directory for discovering and installing Model Context Protocol servers.

AllMCPs on GitHub (opens in a new tab)
Launched onTiny Startupstinystartups.com
Explore
  • Browse servers
  • Best MCP servers
  • Categories
  • MCP clients
  • Agent prompts
  • Stack Builder
  • Compare servers
  • Random discovery New
  • Submit a server
  • Pricing & Boost Boost
Learn
  • Guides hub
  • What is MCP?
  • Install guide
  • Build an MCP server
  • Deploy an MCP server
  • Security guide
  • Troubleshooting
  • MCP for SEO & AEO
  • Protocol versioning
  • Transports: stdio vs HTTP
  • State of MCP (stats)
  • Blog & updates
Tools
  • All developer tools
  • Config generator
  • Config validator
  • Config auditor
  • MCP playground
  • Token calculator
  • OpenAPI β†’ MCP
  • Badge generator
For agents
  • REST API docs
  • Trust & traffic Live
  • Remote MCP server SSE β†— (opens in a new tab)
  • llms.txt β†— (opens in a new tab)
  • Catalog JSON β†— (opens in a new tab)
Company
  • About
  • Advertise Sponsor
  • Contact
  • GitHub β†— (opens in a new tab)
  • Terms
  • Privacy
AllMCPs VerifiedAllMCPs VerifiedFeatured on Nick LaunchesFeatured on Nick LaunchesLaunch Llama NewsletterLaunch Llama NewsletterVerified DR - allmcps.comVerified DR - allmcps.comFeatured on SaaSGrowFeatured on SaaSGrowFeatured on Twelve ToolsFeatured on Twelve ToolsFeatured on Saaspa.geFeatured on Saaspa.geFeatured on Findly.toolsFeatured on Findly.toolsFeatured on Startup FameFeatured on Startup FameFeatured on LaunchKiwiFeatured on LaunchKiwiFeatured on ScrollLaunchFeatured on ScrollLaunchFeatured on DailyPingsFeatured on DailyPingsFazier badgeFazier badgeFeatured on NewTool.siteFeatured on NewTool.siteFeatured on saasfame.comFeatured on saasfame.comDR Checker - Domain RatingDR Checker - Domain RatingListed on Turbo0Listed on Turbo0Launched on LaunchBoard - Product Launch PlatformLaunched on LaunchBoard - Product Launch PlatformList on SimilarlabsList on Similarlabshttps://codetrendy.comhttps://codetrendy.comListed on DevTool.ioFeatured on BuildlistFeatured on BuildlistLaunched on Tiny StartupsFeatured on ShowMeBestAIFeatured on ShowMeBestAIFind us on LaunchZoneFind us on LaunchZoneAllMCPs VerifiedAllMCPs VerifiedFeatured on Nick LaunchesFeatured on Nick LaunchesLaunch Llama NewsletterLaunch Llama NewsletterVerified DR - allmcps.comVerified DR - allmcps.comFeatured on SaaSGrowFeatured on SaaSGrowFeatured on Twelve ToolsFeatured on Twelve ToolsFeatured on Saaspa.geFeatured on Saaspa.geFeatured on Findly.toolsFeatured on Findly.toolsFeatured on Startup FameFeatured on Startup FameFeatured on LaunchKiwiFeatured on LaunchKiwiFeatured on ScrollLaunchFeatured on ScrollLaunchFeatured on DailyPingsFeatured on DailyPingsFazier badgeFazier badgeFeatured on NewTool.siteFeatured on NewTool.siteFeatured on saasfame.comFeatured on saasfame.comDR Checker - Domain RatingDR Checker - Domain RatingListed on Turbo0Listed on Turbo0Launched on LaunchBoard - Product Launch PlatformLaunched on LaunchBoard - Product Launch PlatformList on SimilarlabsList on Similarlabshttps://codetrendy.comhttps://codetrendy.comListed on DevTool.ioFeatured on BuildlistFeatured on BuildlistLaunched on Tiny StartupsFeatured on ShowMeBestAIFeatured on ShowMeBestAIFind us on LaunchZoneFind us on LaunchZone
Β© 2026 Jackalope Digital LLC. All rights reserved.
  1. Home
  2. πŸ’» Developer Tools
  3. TokCalc MCP Server
T
Health: Not checked yetWe have not completed a health check for this listing yet.No health check has run yet.

TokCalc MCP Server

User RatingsBe the first to rate and review this MCP server! Enrichment pendingWe haven’t run our AI enrichment pass on this listing yet, so the overview, use cases, and FAQ below may be sparse or missing. We work through the catalog over time β€” check back soon.
View Repository

LLM capacity planning tools to estimate VRAM, compute, latency, and GPU topology for agents.

Quick Install

Automated & IDE Setup

Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β€” or use 1-click editor setup below.

One-click editor setup isn’t available for this listing yet β€” we don’t have a confirmed install command, and we’d rather show nothing than point your editor at the wrong package or host. Follow the project’s own setup instructions, linked above.

Manual Client & Custom JSON ConfigExpand JSON β–Ύ
No confirmed setup config for this listing yet. We only publish a config block when the install details come from the project itself β€” its README, its docs, or a verified owner. We haven’t found those for TokCalc MCP Server, and we’d rather show nothing than a guess you’d paste into your client. Follow the project’s own setup instructions for the current steps.
Install Directory Badge Claim listing AlternativesπŸ’» More in Developer Tools

Documentation Overview

tokcalc

The open-source LLM serving capacity planner

Plan your LLM deployment before you rent the GPUs.

Live demo License: Apache 2.0 PRs welcome Made with Next.js MCP server


tokcalc demo

Switch model β†’ multi-GPU β†’ long-context capacity planner β†’ Build vs Buy β†’ Reference catalog β†’ Live GPU pricing


tokcalc turns your LLM traffic, context length, latency SLOs, cache behavior, and model choice into a defensible serving topology and cost plan β€” with transparent formulas and cited benchmarks.

It's the engineering-grade pre-deployment decision layer for LLM inference. Not another static "tokens per second" calculator.

Try it now

tokcalc.vercel.app β€” no signup, no tracking, no paywall.

Pick a model, a GPU, and a workload. Get an instant capacity plan:

  • Generation speed (tok/s)
  • Time-to-first-token (TTFT) + inter-token latency (ITL)
  • VRAM budget with KV-cache sizing
  • Multi-GPU topology recommendation (Single GPU β†’ TPΓ—2/4/8 β†’ Context Parallel)
  • Monthly cost + break-even vs API pricing
  • Live GPU pricing across Azure + AWS + GCP + Vast.ai marketplace (fetched in parallel)
  • Shareable URL β€” your config encoded in the URL hash, send to colleagues
  • Copy as Markdown β€” paste the full result into GitHub issues / Slack / docs

Why tokcalc?

The market has dozens of "tokens per second" calculators and self-host-vs-API break-even tools (induwara.lk, gigagpu, kickllm, cloudparity, curlscape, profitable.ai). None of them are unified capacity planners β€” and none show their math.

What tokcalc answers that competitors can't

"Can I serve Qwen 2.5 72B at 128K context on 2Γ— H100 with 20 concurrent users?"

"How many H200s do I need for 1,000 req/min with P95 TTFT < 2s?"

"Does FP8 or AWQ save more money once quality, KV cache, and engine support are included?"

"At what daily volume does an H100 beat GPT-4o pricing?"

"What's the cheapest H100 right now across Azure / AWS / GCP / Vast.ai?"

"What happens to cost and latency if an agent makes 8 model calls, has 3 tool calls, and its context grows by 5K tokens each turn?"

"Would prefix caching, continuous batching, or PD disaggregation save more for this workload?"

Features

Calculator tab

FeatureWhat it computes
Model fit / VRAMWill the model + KV cache fit in the GPU's memory?
ThroughputDecode tok/s (per-stream) + aggregate (batched) + prefill tok/s
Latency splitTime-to-first-token (= prefill) + inter-token latency (= decode)
Continuous batchingUser-tunable 1.0–4Γ— multiplier (cited 1.5–4Γ— SOSP range)
Reasoning tokensHidden reasoning budget added to billed output (o1 / R1 / Claude thinking)
Prompt cachingSelf-hosted vLLM APC + Anthropic 5m/1h TTL + OpenAI 50%-off cached tokens
Speculative decodingUser-tunable 1.2–4Γ— boost factor
Multi-GPU TP1Γ— β†’ 8Γ— tensor parallel with NVLink efficiency factor
Long-context capacityKV memory + max concurrency + prefill time at 4K β†’ 1M context
Topology recommendationSingle GPU β†’ TPΓ—2 β†’ TPΓ—4 β†’ TPΓ—8 β†’ TPΓ—8 + Context Parallel (RingAttention)
Cost economicsGPU $/hr β†’ $/M output tokens β†’ $/request β†’ monthly cost
Observed benchmark calibrationPaste vLLM/SGLang/TRT-LLM JSON β†’ see formula accuracy verdict (validated / underestimated / overestimated)

Build vs Buy tab

Independent calculator (separate state) that compares:

  • Self-host: model + GPU + quant + utilization + batch β†’ $/M tokens + monthly cost
  • API: 13 providers (OpenAI / Anthropic / Gemini / Groq / DeepSeek / Mistral / Together)
  • Verdict: Self-host cheaper / API cheaper / Not enough volume β€” with break-even reqs/day

Reference tab

6 sub-tables β€” fully transparent, every record source-linked where available:

  • Models (35 entries: Llama 4 Scout/Maverick, Qwen 3 family, DeepSeek V3/R1, Pixtral, BGE-M3, ...)
  • GPUs (30 entries: H100/H200/B200/B300, AMD MI300X/MI325X, Intel Gaudi 3, TPU v5p/Trillium, Groq LPU, Cerebras WSE-3, Apple M2/M3/M4 Ultra, ...)
  • Quantization (16 formats: FP16/BF16, GGUF Q2_Kβ†’Q8_0, GPTQ, AWQ, EXL2, FP8, NVFP4)
  • API pricing (13 models with input/cached/output + retired/current status)
  • Cloud GPU pricing (all GPUs with $/hr > 0 + typical providers β€” static curated)
  • πŸ”΄ Live pricing (real-time fetch from Azure + AWS + GCP + Vast.ai in parallel β€” LIVE badges + timestamps)

The long-context capacity planner (the differentiator)

This is the formula the Perplexity research brief called "the most important tokcalc should visibly expose":

$$ \text{KV bytes/request} = 2 \cdot L \cdot T \cdot H_{\text{kv}} \cdot D_h \cdot B $$

For dense attention, prefill cost grows superlinearly with context length:

$$ \text{prefill FLOPs} = \underbrace{2 \cdot N \cdot T}{\text{linear}} + \underbrace{T^2 \cdot H{\text{kv}} \cdot D_h \cdot L}_{\text{attention}} $$

tokcalc shows you:

  • Max concurrent users at 4K / 8K / 16K / 32K / 64K / 128K / 256K / 512K / 1M context
  • KV memory per request at each context length
  • Prefill time (with superlinear attention correction beyond 32K)
  • Required topology (Single GPU β†’ TPΓ—2/4/8 β†’ TPΓ—8 + Context Parallel)
  • RingAttention citation when CP is needed

Live GPU pricing (NEW)

The πŸ”΄ Live pricing sub-tab in Reference fetches real-time GPU prices from 4 providers in parallel via Promise.allSettled:

ProviderAPIAuthCache TTL
AzureRetail Prices APINone (public)24h
AWSEC2 bulk pricing fileNone (public)24h
GCPCloud Billing Catalog APIGCP_API_KEY env var24h
Vast.aiMarketplace bundles APINone (public)5m (spot prices change rapidly)

Each provider shows a LIVE badge with timestamp + cache state. The comparison table shows the cheapest price per GPU across all 4 providers (highlighted in emerald) + per-provider breakdown.

If any provider fails (e.g., GCP_API_KEY not set), the others still work β€” graceful degradation per provider.

The math, transparently

Every number above comes from a formula you can inspect. No black boxes.

Decode tokens/sec (memory-bandwidth bound)

$$ \text{decode tok/sec} \approx \frac{\text{HBM BW} \cdot \eta_{\text{mem}} \cdot \text{quant_eff}}{\text{model size}} $$

Where:

  • HBM BW = GPU memory bandwidth (e.g., 3350 GB/s for H100 SXM)
  • Ξ·_mem = 0.65 = typical real-world memory utilization (35% overhead)
  • quant_eff = dequantization efficiency multiplier (1.0 for FP16, 1.5 for FP8 on H100, 0.85 for INT4)
  • model size = active_params Γ— bytes_per_param (uses ACTIVE params for MoE, not total)

Refs: PagedAttention paper (arxiv.org/abs/2309.06180)

Prefill tokens/sec (compute bound)

$$ \text{prefill tok/sec} \approx \frac{\text{GPU FLOPS} \cdot \eta_{\text{compute}}}{2 \cdot \text{active params}} $$

Where:

  • GPU FLOPS = dense FP16/BF16 TFLOPS (sparse values not used)
  • Ξ·_compute = 0.50 = typical compute utilization
  • Factor of 2 = one multiply + one add per parameter per token

For long context (>32K), the superlinear attention correction above applies.

Continuous batching multiplier (workload-specific)

$$ \text{aggregate tok/sec} = \text{decode tok/sec} \cdot \text{batch size} \cdot \text{continuous batching multiplier} $$

Critical caveat: There is no universal continuous batching multiplier. vLLM reported 14–24Γ— vs HF Transformers (extreme), 2.2–2.5Γ— vs TGI. SOSP paper finds 2–4Γ— typical vs FasterTransformer/Orca. tokcalc defaults to a conservative 1.5Γ— and lets you tune.

Refs:

  • vLLM blog (2023-06-20)
  • PagedAttention paper
  • Anyscale continuous batching study
Prompt caching economics (Anthropic / OpenAI)

For shared prefix of length $T_p$, suffix of length $T_u$, output $O$, cache hit rate $h$:

$$ \text{API input cost} = N \cdot \left[ (1-h) \cdot T_p \cdot P_{\text{write}} + h \cdot T_p \cdot P_{\text{read}} + T_u \cdot P_{\text{input}} \right] $$

Anthropic multipliers (verified 2025-2026):

  • 5-minute cache write: 1.25Γ— base input
  • 1-hour cache write: 2.0Γ— base input
  • Cache read: 0.1Γ— base input (90% savings)

OpenAI: cached input discounted 50%, no separate write fee.

Refs:

  • Anthropic prompt caching docs
  • OpenAI prompt caching guide
Multi-GPU topology recommendation

$$ \text{total needed} = \text{model weights} + (\text{KV per request} \cdot \text{batch size}) $$

Walk the smallest topology that fits:

  • Single GPU: total ≀ VRAM Γ— 1
  • Tensor Parallel Γ—2/4/8: total ≀ VRAM Γ— N (weights + KV split evenly)
  • TPΓ—8 + Context Parallel: total > VRAM Γ— 8 β€” use RingAttention to shard KV across nodes

Refs: RingAttention paper

Self-host vs API break-even

Read the full README β†’View source on GitHub β†’

Related MCP Servers

View all in Developer Tools View all alternatives
  • O
    Openapi MCP Server

    Connect any HTTP/REST API server using an Open API spec (v3)

    πŸ’» Developer Tools3 views
    Compare vs Openapi MCP Server β†’
  • C
    Claude Task Master

    AI-powered task management system for AI-driven development. Features PRD parsing, task expansion, multi-provider support (Claude, OpenAI, Gemini, Perplexity, xAI), and selective tool loading for optimized context usage.

    πŸ’» Developer Tools8 views
    Compare vs Claude Task Master β†’
  • M
    MCP Server Docker

    Integrate with Docker to manage containers, images, volumes, and networks.

    πŸ’» Developer Tools3 views
    Compare vs MCP Server Docker β†’
  • N
    Next Devtools MCP
    Verified

    Official Next.js MCP server for coding agents. Provides runtime diagnostics, route inspection, dev server logs, docs search, and upgrade guides. Requires Next.js 16+ dev server for full runtime features.

    πŸ’» Developer Tools6 views
    Compare vs Next Devtools MCP β†’

Reviews

No reviews yet β€” be the first to share how this listing worked for you.

Frequently Asked Questions about TokCalc MCP Server

We don't have a confirmed install command for TokCalc MCP Server yet, so we don't publish a generated one β€” a guessed package name would point at the wrong package or none at all. Follow the project's own README or setup instructions (https://github.com/stevecrates489-commits/tokcalc.git) for the current steps.

AllMCPs Directory Badge

Full Badge Customizer

Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.

Badge Style:
Live Dynamic SVG PreviewTokCalc MCP Server AllMCPs Directory Badge
Markdown (GitHub README)
[![AllMCPs](https://allmcps.com/api/badge/tokcalc-mcp-server?style=directory)](https://allmcps.com/mcp/tokcalc-mcp-server)
HTML Embed
<a href="https://allmcps.com/mcp/tokcalc-mcp-server"><img src="https://allmcps.com/api/badge/tokcalc-mcp-server?style=directory" alt="TokCalc MCP Server on AllMCPs" /></a>

Technical Specs & Signals

CategoryπŸ’»Developer Tools
More technical detailsExpand β–Ύ
Last updatedSep 28, 2026
Views0
Unique ViewsTotal visits recorded for this listing page on AllMCPs.
Installs0
Installs & Copy ActionsTotal times users copied install commands or configuration snippets for this server.
27Quality signal: Emerging Β· 27/100How this signal is calculated β–Ύ
Server availabilityNot measured

Not scored for repo-hosted servers β€” we can't reach the running server, only its GitHub page. Hosted MCP endpoints are health-checked live.

Verified ownership8/20
Documentation & tools11/30
Adoption & activity1/15
Community engagement0/10

A guidance signal from public completeness & health data β€” not a user rating. New listings start lower and rise as they add docs, get verified, and grow adoption. Signals we can't observe for a listing are skipped, not counted against it.

β˜… Spotlight Slot

Feature Your MCP Server

Get maximum visibility for your server across our directory, search results, and detail pages.

Spotlight Your Server

Own this project?

This directory is pre-filled from public sources. Claim via GitHub README, site badge, or DNS TXT to unlock edit access and the Official badge and attach your website β€” proof is checked automatically, then reviewed by our team.

Free dofollow backlink: add your website and place the AllMCPs badge on it β€” no claim needed. We detect it automatically and keep it verified as long as the badge stays live.

Claim & get free dofollow

Share & Embed

Add our SVG badge (dark/light directory styles) or embeddable widget to your site.

Explore more

More in πŸ’» Developer Tools β†’Best MCP servers for Developers β†’Alternatives to TokCalc MCP Server β†’Install in Claude DesktopInstall in CursorInstall in VS CodeSetup guides for all 13 MCP clients