Design transformer LLM architectures and report their parameters, FLOPs, memory and cost
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent — or use 1-click editor setup below.
One-click editor setup isn’t available for this listing yet — we don’t have a confirmed install command, and we’d rather show nothing than point your editor at the wrong package or host. Follow the project’s own setup instructions, linked above.
Schematic capture for neural network architectures.
Draw the model, get the numbers, generate the PyTorch.
Open the editor · Quick start · What it does · Correctness · For agents · Docs · Contributing · Roadmap · License
TensorCAD treats a neural network the way an EDA tool treats a circuit. Blocks are symbols with typed pins. Tensors are nets. Shapes are checked by a real algebra rather than by running the thing. A design-rule check tells you the model will not fit on your GPUs before you rent them. And the drawing is not a picture of the model — it is the model, and PyTorch falls out of it.
It started as a tool for language models. It now draws vision transformers and convolutional classifiers too, because the same machinery turned out to work.
Llama-3-8B, one level open. Every net carries its shape; the readout is measured under the operating point at the top of it, and says how the count compares with the published one.
Fit it on the cluster. Every split the cluster admits, priced and ordered by how little it asks of you. Pressing one applies it, and the whole readout follows. |
Sweep small, run big. The same design at several widths, with what to multiply the initialization and the learning rate by at each. Press a rung to open it. |
The same model as volumes, a port of Brendan Bycroft's LLM visualisation (MIT). Blue is a weight, green an activation, and a ribbon is a tensor on its way somewhere.
Every picture here is regenerated by bun run scripts/screenshots.ts from a
running dev server, so it is what the tool looks like rather than what it
looked like once. bun run scripts/export-svg.ts writes the sheet itself as a
vector — one is in docs/images — which is
also what File > Export the sheet as SVG does from the editor.
It runs in a browser: tensorcad.dev. There is no
server behind it — the engine is the same WebAssembly module the command line
and the MCP server load, so every number on the screen is computed in the tab,
and this build has nowhere to send a design even if it wanted to: it registers
no storage provider, so File > Save a copy writes to your disk and that is
the only copy there is.
A separate deployment at app.tensorcad.dev is this same editor with an account attached, where designs are saved and can be shared. The first claim holds there too — the analysis still runs in the tab — but the second does not, which is why they are two sentences and two addresses rather than one of each.
The documentation is at docs.tensorcad.dev.
Needs Bun and Go 1.25 or later. The analysis engine is Go compiled to WebAssembly, which is the one build step:
Every number in the project comes out of one command:
Generate a model and check it against real PyTorch:
Open the editor:
Both load the same engine. So do the command line and the MCP server, which is the point: the numbers cannot depend on where you asked for them.
Schematic capture. Blocks, typed pins, orthogonal wire routing, four-sided
pin anchors, junction dots on branching nets, hollow circles on unconnected
pins. Containers unfold in place so a 32-layer stack reads as one frame with a
32× bracket, the way published architecture figures draw it.
A feature timeline. Every edit is kept as an operation with its arguments, not as a snapshot, so a step can be taken out of the middle and everything after it replays on top of what is left. Suppress the step that widened the model and the rename you did afterwards survives. A step that cannot replay — because you suppressed the one that added the block it wired — says what it could not find rather than being dropped in silence.
Tensors you can point at. A wire is a tensor, and clicking one says what it
carries: its shape, its dtype, the block that made it, every block that reads
it, and its share of the activation memory. Every segment of the same net
lights with it. A block that fans out — Nemotron-H's split holds 290 MiB
across three output pins, 128, 160 and 2 — is where that matters: its own
number answers neither which of them is the big one nor what dropping one would
save.
A real shape algebra. Every tensor dimension is a multivariate polynomial
with exact rational coefficients over named symbols. B and T stay
indeterminate all the way through, so a mismatch is a genuine polynomial
difference rather than two numbers that happened not to match. Splits carry
divisibility obligations instead of silently rounding.
Design-rule checks. Eighteen rules: head divisibility, vocabulary padding, RoPE dimension parity, interface breakage, whether the design fits the GPUs you selected under the sharding plan you chose. The DRC panel correctly refuses Llama-3-8B at 90.16 GiB/GPU against an H100's 80.
Quantitative analysis. Parameters, FLOPs, activation memory (Megatron
formulas), KV cache, ZeRO/FSDP/TP/PP sharding, roofline throughput, Chinchilla
budgets. Nothing in the UI computes its own numbers; one validate() call per
document and operating point feeds every panel.
PyTorch generation. generateTorch(doc) emits a runnable model with an
init_weights() method — because nn.Embedding defaults to a unit normal, and
that is the difference between a next-token loss of 466 and 10.94 against the
ln(50257) = 10.82 baseline.
A 3D volume view. Every tensor as a plate, sized by its real dimensions, with flow ribbons between them. Ported from Brendan Bycroft's LLM visualisation.
An MCP server. So an agent can design, validate, analyse and generate without a human in the loop.
This is the part worth reading, and the reason to trust the numbers.
Twenty presets are the regression suite. Each one carries the parameter count its authors published, and the tests assert the analysis reproduces it. Seventeen match to the parameter; the other three are checked against rounded vendor figures with an explicit tolerance.
Every preset is instantiated in real PyTorch. python -m tensorcad_runtime verify builds the generated model on the meta device and reports its true
parameter count, module by module, up to DeepSeek-V3 at 671,026,419,200.
No reviews yet — be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/tensorcad)<a href="https://allmcps.com/mcp/tensorcad"><img src="https://allmcps.com/api/badge/tensorcad?style=directory" alt="Tensorcad on AllMCPs" /></a>