The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Tensorcad listing page.
Schematic capture for neural network architectures.
Draw the model, get the numbers, generate the PyTorch.
Open the editor · Quick start · What it does · Correctness · For agents · Docs · Contributing · Roadmap · License
TensorCAD treats a neural network the way an EDA tool treats a circuit. Blocks are symbols with typed pins. Tensors are nets. Shapes are checked by a real algebra rather than by running the thing. A design-rule check tells you the model will not fit on your GPUs before you rent them. And the drawing is not a picture of the model — it is the model, and PyTorch falls out of it.
It started as a tool for language models. It now draws vision transformers and convolutional classifiers too, because the same machinery turned out to work.
Llama-3-8B, one level open. Every net carries its shape; the readout is measured under the operating point at the top of it, and says how the count compares with the published one.
Fit it on the cluster. Every split the cluster admits, priced and ordered by how little it asks of you. Pressing one applies it, and the whole readout follows. |
Sweep small, run big. The same design at several widths, with what to multiply the initialization and the learning rate by at each. Press a rung to open it. |
The same model as volumes, a port of Brendan Bycroft's LLM visualisation (MIT). Blue is a weight, green an activation, and a ribbon is a tensor on its way somewhere.
Every picture here is regenerated by bun run scripts/screenshots.ts from a
running dev server, so it is what the tool looks like rather than what it
looked like once. bun run scripts/export-svg.ts writes the sheet itself as a
vector — one is in docs/images — which is
also what File > Export the sheet as SVG does from the editor.
It runs in a browser: tensorcad.dev. There is no
server behind it — the engine is the same WebAssembly module the command line
and the MCP server load, so every number on the screen is computed in the tab,
and this build has nowhere to send a design even if it wanted to: it registers
no storage provider, so File > Save a copy writes to your disk and that is
the only copy there is.
A separate deployment at app.tensorcad.dev is this same editor with an account attached, where designs are saved and can be shared. The first claim holds there too — the analysis still runs in the tab — but the second does not, which is why they are two sentences and two addresses rather than one of each.
The documentation is at docs.tensorcad.dev.
Needs Bun and Go 1.25 or later. The analysis engine is Go compiled to WebAssembly, which is the one build step:
Every number in the project comes out of one command:
Generate a model and check it against real PyTorch:
Open the editor:
Both load the same engine. So do the command line and the MCP server, which is the point: the numbers cannot depend on where you asked for them.
Schematic capture. Blocks, typed pins, orthogonal wire routing, four-sided
pin anchors, junction dots on branching nets, hollow circles on unconnected
pins. Containers unfold in place so a 32-layer stack reads as one frame with a
32× bracket, the way published architecture figures draw it.
A feature timeline. Every edit is kept as an operation with its arguments, not as a snapshot, so a step can be taken out of the middle and everything after it replays on top of what is left. Suppress the step that widened the model and the rename you did afterwards survives. A step that cannot replay — because you suppressed the one that added the block it wired — says what it could not find rather than being dropped in silence.
Tensors you can point at. A wire is a tensor, and clicking one says what it
carries: its shape, its dtype, the block that made it, every block that reads
it, and its share of the activation memory. Every segment of the same net
lights with it. A block that fans out — Nemotron-H's split holds 290 MiB
across three output pins, 128, 160 and 2 — is where that matters: its own
number answers neither which of them is the big one nor what dropping one would
save.
A real shape algebra. Every tensor dimension is a multivariate polynomial
with exact rational coefficients over named symbols. B and T stay
indeterminate all the way through, so a mismatch is a genuine polynomial
difference rather than two numbers that happened not to match. Splits carry
divisibility obligations instead of silently rounding.
Design-rule checks. Eighteen rules: head divisibility, vocabulary padding, RoPE dimension parity, interface breakage, whether the design fits the GPUs you selected under the sharding plan you chose. The DRC panel correctly refuses Llama-3-8B at 90.16 GiB/GPU against an H100's 80.
Quantitative analysis. Parameters, FLOPs, activation memory (Megatron
formulas), KV cache, ZeRO/FSDP/TP/PP sharding, roofline throughput, Chinchilla
budgets. Nothing in the UI computes its own numbers; one validate() call per
document and operating point feeds every panel.
PyTorch generation. generateTorch(doc) emits a runnable model with an
init_weights() method — because nn.Embedding defaults to a unit normal, and
that is the difference between a next-token loss of 466 and 10.94 against the
ln(50257) = 10.82 baseline.
A 3D volume view. Every tensor as a plate, sized by its real dimensions, with flow ribbons between them. Ported from Brendan Bycroft's LLM visualisation.
An MCP server. So an agent can design, validate, analyse and generate without a human in the loop.
This is the part worth reading, and the reason to trust the numbers.
Twenty presets are the regression suite. Each one carries the parameter count its authors published, and the tests assert the analysis reproduces it. Seventeen match to the parameter; the other three are checked against rounded vendor figures with an explicit tolerance.
Every preset is instantiated in real PyTorch. python -m tensorcad_runtime verify builds the generated model on the meta device and reports its true
parameter count, module by module, up to DeepSeek-V3 at 671,026,419,200.
FLOPs are checked against a profiler. For AlexNet the agreement is exact —
1,428,376,960 per image, ratio 1.000000 against torch.utils.flop_counter. For
GPT-2 small the profiler says 251.78 MFLOP/token and the analysis says 249.42,
and the whole difference is the causal mask: a profiler counts the attention
operator as if nothing were masked. flops.fwdTotalUnmasked reproduces the
profiler exactly; flops.fwdTotal is what a fused causal kernel actually does.
A test pins both numbers.
The Go port is proven against the TypeScript it replaces. The engine is migrating to Go; the TypeScript writes golden files for all twenty presets and the Go tests must reproduce them exactly — including the evaluation order of the symbol table, the printed form of every polynomial, and the text of every error.
| Language | GPT-2 (small→XL), nanoGPT, Llama 2/3/3.1, Mistral, Qwen 2.5/3, Gemma 2, Mixtral, DeepSeek-V3, Nemotron-H |
| Vision | I-JEPA ViT-H/14 — bidirectional attention, three towers including the EMA target encoder |
| Convolutional | AlexNet — B C H W tensors, spatial downsampling, 96% of its weights in the classifier |
Mechanisms covered: GQA/MQA/MHA, multi-head latent attention, SwiGLU/GeGLU, RMSNorm/LayerNorm, RoPE with scaling, mixture-of-experts with shared experts and routing bias, Mamba-2 state-space layers, sliding-window attention, QK-norm, post-norm, tied embeddings, hybrid stacks.
One engine, everywhere. The editor, the command line, the MCP server and the
desktop shell all load the same WebAssembly module and ask it the same
questions, so an answer cannot depend on where it was asked. What it is held to
is packages/core-go/testdata: every preset's symbol table, inferred shapes,
full analysis, design-rule findings and generated PyTorch, byte for byte,
checked both against the Go source and against the compiled module. Those files
began as the answers of the TypeScript this was ported from, which has since
been deleted.
This repository is written to be worked on by coding agents as well as people.
CLAUDE.md is the entry point: commands, layout, the eight
invariants, and what to do when adding a block. Read it first.B and T are reserved. Break
one and the tests will tell you, but the design rules will not.docs.summary and
docs.formula with a source link, and — for a primitive — paramCount,
flops, retains and stateBytes. Then a preset that uses it with a
published figure, or a test pinning the arithmetic. Run
bun run scripts/report.ts before and after..mcp.json.TENSORCAD_BRIDGE=1 and the editor attaches to it. The
agent's edits appear on the canvas as it makes them and land on the undo
stack, so a person watching can take one back; what that person does comes
back the other way. There is one document, not two — the editor's edits go
through the same revision-checked apply a tool call does. It binds
127.0.0.1, refuses a foreign Origin, and opens no port at all without the
variable. See
Drive TensorCAD from an agent.TensorCAD borrows from work that deserves naming:
Formulas are sourced individually in
docs/reference/analysis-math.md,
including two figures the original research got wrong that the implementation
corrects.
Working and useful, with rough edges. The Go migration is at stage 2 of 10. See
ROADMAP.md for what is known to be missing — linear-attention
blocks, multi-token prediction, and Gemma's alternating local/global attention,
which the importer warns about rather than approximating.
docs/, organised by Diátaxis:
tutorials to learn from,
how-to guides to work from,
reference to look things up in, and
explanation for why any of it is
the way it is.
MIT. See LICENSE.md, which also carries the notices for the
MIT-licensed work this project ports.