The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the FitLLM listing page.
FitLLM is an open-source, zero-dependency engine that checks whether a local LLM fits on a GPU or Apple Silicon Mac using architecture-aware memory math.

Live: https://fitllm.run · Bilingual · Free · No ads · No login
Open engine: fitllm-engine (MIT · npm
fitllm-engine·npx fitllm)Zero dependencies. One readable file:
engine.js. Conformance-vector tested. MIT.
Connect any Streamable HTTP MCP client to https://fitllm.run/api/mcp:
check_llm_fit — check one model against a GPU, multi-GPU rig, or Mac and return the verdict, memory breakdown, and a fix when it does not fit.what_fits_on_hardware — rank the supported local models that fit the given GPU, multi-GPU rig, or Mac.list_supported — list the built-in model and hardware names accepted by the fit checker.The server is read-only, stateless, and requires no authentication.
--detect reads all nvidia-smi adapters, uses Apple Silicon unified memory on arm64 macOS, and can resolve an exact catalog GPU name through Windows/WSL PowerShell. It never uses Win32_VideoController.AdapterRAM, serials, PNP IDs, or full environment dumps. Intel-only, ambiguous, and unsupported adapters stop with exit 2 instead of borrowing a nearby GPU's memory.
Why a CLI? The "will it run?" question is born in the terminal — one line before ollama pull. No install, no tab-switching, and it reads your actual hardware with --detect instead of asking you to know your VRAM. Exit code 0/1 makes it a pre-download guard:
The CLI exit contract composes directly with model runtimes. The download or launch runs only after a FITS or TIGHT verdict; invalid inputs stop with exit 2.
For CI, use the repository's composite Action. It needs no secret and returns the full --json --why result as steps.preflight.outputs.result:
Set exactly one of gpu or mac. The Action preserves the CLI contract: exit 0 means fits/tight, 1 means it will not fit, and 2 means the request is invalid. Inputs cross the GitHub expression boundary through environment variables and are passed to Node as a Bash argument array.
Agents can collect a typed measurement without uploading anything:
The command validates the conditions and prints a candidate JSON object plus a prefilled GitHub issue URL. Submission remains a human action. Public Hugging Face IDs (org/model) are fetched with a bounded config/index reader and accepted only when parseHfConfig() supports the architecture.
This is the open calculation core of FitLLM. The math is open so you can audit it.
Ask an LLM "does Qwen 3.6 fit my GPU?" and it pattern-matches to an architecture from its training cutoff — and usually says no. Catalog-based calculators lag new releases. The CLI, API, and MCP use a curated catalog pinned to official configs. The web calculator can additionally inspect a pasted Hugging Face ID's official config.json live, so supported architectures work on day-one releases — including the hybrid / sliding-window / MoE structures that naive formulas get wrong.
Covers Apple Silicon unified memory (M1–M6, Pro/Max/Ultra — up to the 512GB Mac Studio), NVIDIA GPUs (RTX 20/30/40/50, workstation RTX 6000 Ada / RTX PRO 6000, datacenter A100/H100/H200/B200), AMD Radeon (RX 7000/9000, PRO W7900) and multi-GPU presets (2×3090, 2×4090, 4×3090) — with GGUF Q-tier weight quantization kept separate from KV-cache quantization. Hardware entries carry their source URLs per-value in engine.js; new entries require ≥2 independent sources (CONTRIBUTING).
Almost every "can I run this LLM?" calculator estimates the KV cache with the textbook formula:
That assumes every layer keeps a full-context KV cache with one uniform head shape. True for Llama-1/2 — wrong for most 2025–2026 models:
| Model | What naive formulas miss | Naive KV | FitLLM KV | Off by |
|---|---|---|---|---|
| Gemma 4 31B @131K, 8-bit | 50 of 60 layers are sliding-window (keep only the last 1024 tokens); the 10 global layers use a different head shape (4 KV-heads × 512, not 16 × 256) | ~60 GB | ~5.4 GB | 11× |
| Qwen 3.6 27B @131K, 8-bit | 48 of 64 layers are linear attention (Gated DeltaNet) — no growing KV cache | ~16 GB | ~4 GB | 4× |
| Qwen 3.8 27B @256K, F16 KV | same shape, newest generation: KV lives on 16 of 64 layers only | 64.0 GiB | 16.0 GiB | 4× |
| GLM-4.7-Flash @128K, bf16 | MLA: K/V compressed into one shared latent (512+64 dims, cached once — not per-head K and V) | ~117 GB | ~6.6 GB | 17.8× |
| Plain dense (Llama, Mistral…) | nothing — standard transformer | same | same | 1× ✅ |
An 11× error flips the verdict: a naive calculator says Gemma 4 31B won't fit in 64 GB at long context, when it fits comfortably.
linearState) — it is a constant per sequence, so it never inflates the context curve.kv_lora_rank + RoPE dims) shared across all heads — per-head "2 × heads × head_dim" formulas over-count by an order of magnitude. Verified against the DeepSeek-V2 paper (arXiv:2405.04434) and the official DeepSeek-V3 inference code.head_dim (Gemma 4: 512 vs 256). MoE keeps every expert in memory while activating only a few per token.per_layer_token_embd input-layer tensor to CPU/host buffers instead of accelerator memory, so only the non-PLE weights are counted against VRAM for the verified Gemma 4 e2b/e4b entries. Counting all 5.1B params against a GPU over-predicts e2b's resident weights by ~1.9× and flips small-card verdicts. This deduction is conditional and rests on that placement alone: the host memory the tensor needs is not budgeted by the GPU verdict, a runtime that loads PLE tensors onto the accelerator (vLLM, for example) invalidates the estimate, and unverified families keep their full weights resident. On Apple Silicon unified memory the full weights stay counted. The exact premise and its pinned sources are listed under Structural premises; a direct measurement on a Gemma 4 GGUF is welcome in issue #7.This engine models each layer type separately, verified against official HuggingFace config.json files.
Plus a parseHfConfig() that turns configs from verified, modeled Hugging Face architecture families into the model shape above; unsupported structures fail closed instead of returning a guess. (No token/s prediction — deliberately: speed depends on runtime/backend in ways a static model can't claim honestly. Fit is a verifiable claim; speed is not.)
Some verdicts are valid only under a structural premise about the artifact or the runtime path. The engine derives the active premises from the same normalized model fields the calculation uses and attaches them to affected results as structuralAssumptions, an array of { id, statement }. Unaffected results, such as plain GQA models, carry no such key, so legacy JSON shapes are unchanged. The CLI prints each active premise as a premise [id]: line in text mode and includes the array in --json (also per row in --top --json); --why and the composite Action forward the same array. There is no runtime selector, no confidence score, and no speed claim: a premise tells you what the memory math assumes, not how fast the model runs.
mla-compressed-latent-cache — KV memory assumes a compressed-latent MLA artifact or mode; legacy non-MLA GGUF or an explicitly uncompressed mode invalidates this estimate. Sources: DeepSeek-V3 inference model.py · llama.cpp convert_hf_to_gguf.py · llama.cpp llama-kv-cache.cpp.ple-llamacpp-non-gpu-residency — GPU weight memory excludes the verified Gemma 4 PLE tensors only because the pinned llama.cpp/GGUF path assigns the per_layer_token_embd input-layer tensor to CPU/host buffers instead of accelerator memory; that host memory is not budgeted here, and a runtime that loads PLE onto the accelerator invalidates this estimate. Sources: Gemma 4 E2B config.json (text body model_type is exactly gemma4_text) · pinned llama.cpp 8b4b3558f1459c13e4aa38d5c94d306a00dc6acd: input-layer device assignment · PER_LAYER_TOKEN_EMBD as an input-layer tensor · Gemma 4 construction · loader header · loader implementation. The deduction applies only to the catalog entries Gemma 4 e2b and Gemma 4 e4b (pleOffloadVerified: true) and to parsed configs whose text body is exactly gemma4_text and whose full profile (layer count, hidden_size, intermediate_size and both PLE dimensions) matches a pinned official E2B/E4B config with a checkpoint that reconciles against the config's own dense body. A look-alike field on any other family, an altered dimension or layer count, or an inconsistent checkpoint keeps the full weights resident on the GPU, so an unverified config cannot flip a verdict toward "fits".mtp-ordinary-generation — KV memory assumes ordinary non-speculative generation; an MTP draft context is not included. Sources: llama.cpp llama-hparams.cpp · llama.cpp llama-model.cpp · llama.cpp convert_hf_to_gguf.py · vLLM qwen3_next_mtp.py.Resident scope. SSD/NVMe streaming, expert paging, swap and general CPU/RAM offload stay outside ordinary FitLLM fit verdicts. The one disclosed structural exception is a placement fact, not storage paging: for the verified Gemma 4 e2b/e4b entries the discrete-GPU verdict excludes the per-layer-embedding (PLE) tensor because the pinned llama.cpp/GGUF path deterministically assigns that input-layer tensor to CPU/host buffers instead of accelerator memory (premise ple-llamacpp-non-gpu-residency, shown with the verdict). That host/system memory is not budgeted by the discrete-GPU verdict, a runtime that loads PLE onto the accelerator invalidates the PLE estimate, and on Apple unified memory the full weights stay counted.
From a repository checkout, recompute the typed measured-vs-predicted ledger with npm run benchmark:accuracy (or add -- --json). Run the pinned estimator differential with npm run benchmark:differential. To recapture its uncut source output from a separately downloaded official llmfit release artifact:
The differential is labeled architecture_differential_not_runtime_accuracy: it demonstrates how estimators treat GQA, sliding-window attention, and MLA, but it is not runtime measurement evidence. The methodology, competitor input format, checksums, and precommitted kill conditions are in benchmarks/README.md. The current evidence does not clear the comparative accuracy claim gate.
config.json.All figures are estimates — real usage varies with the runtime (MLX/Ollama/llama.cpp), OS state, and quantization scheme.
vectors/fit-vectors-v1.json pins 30 language-neutral test vectors (exact KV bytes, per-token costs, fit verdicts) derived by hand from official config.json values — e.g. "Gemma 4 31B at 262,144 ctx, bf16 = exactly 22,313,697,280 bytes". Any implementation in any language conforms if every vector passes — run ours with node vectors/run.mjs.
Why this matters: the formulas are easy to copy; a verified answer key is not. If you port this engine to Python, Rust or Go, you don't become an untrusted fork — pass the vectors and you're a conformant implementation of the same standard. Port the engine, keep the vectors.
census/ holds 9,477 verdicts (27 models incl. draft tier × 93 GPUs/Macs × quant tiers) computed by this engine — as CSV/JSON you can import, chart or cite, plus a starter matrix ("biggest model that fits comfortably per device"). Regenerate it yourself: npm run census. Real-world measurements land next to predictions via fixtures/ PRs — predicted vs. measured, in public.
Show whether a model runs on given hardware — live from the engine, one line in any README or model card:
Params: model (name, fuzzy), gpu (name, fuzzy) or ram (GB, Apple unified memory), optional quant (GGUF tier / 4|8|16), ctx, kv. Verdict color: green fits · yellow tight · red won't fit.
Why embed it? The #1 question under every model card and local-AI tutorial is "will it run on my machine?" The badge answers it live from the engine — recomputed when the data updates, not a stale claim frozen into your README. If you publish models or write guides: one line replaces a whole FAQ paragraph and cuts the "it OOM'd on my 8GB card" issues before they're filed.
The engine runs as a public MCP server at https://fitllm.run/api/mcp — connect it once and your assistant answers "can I run X on my Y?" with this engine's math instead of guessing from stale training data (LLMs routinely get KV-cache math wrong — see the 17.8× table above).
https://fitllm.run/api/mcpclaude mcp add --transport http fitllm https://fitllm.run/api/mcpmcp.json → { "mcpServers": { "fitllm": { "url": "https://fitllm.run/api/mcp" } } }Tools: check_llm_fit (verdict + full memory breakdown + fix suggestion — supports multi-GPU rigs like "RTX 5090 + RTX 3090"), what_fits_on_hardware (ranked list for your machine), list_supported. Resources: fitllm://models, fitllm://hardware, fitllm://census, fitllm://engine. Intentionally open: read-only, stateless, no auth, no secrets — every call is a pure function of public data.
Listed on: official MCP registry (run.fitllm/fitllm) · Glama · mcp.so · Smithery
No MCP client? One GET, no auth, no key — JSON by default, plain text for curl:
Open data: the full Fit Census (9,477 verdicts, CC0) at fitllm.run/data and on Hugging Face Datasets. Try the engine in-browser: HF Space demo.
No ads. No login. No affiliate links. Output is never for sale. Fit is a winnable, verifiable claim; raw tok/s is not — so this engine refuses speed predictions rather than dress a guess as precision.
Ran a model and measured real peak memory? Report a measurement — it improves the estimates for everyone.
yongha — GitHub. Powers fitllm.run.
MIT © click6067-ship-it