Code changes and unit tests by local MLX models, behind a gate a machine can run.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
One-click editor setup isnβt available for this listing yet β we donβt have a confirmed install command, and weβd rather show nothing than point your editor at the wrong package or host. Follow the projectβs own setup instructions, linked above.
You ask for a change to your source code, and you get it β made on your Mac, by models that cost you no Claude tokens, and only after a machine has checked it.
Claude Opus plans the work and reviews what came back. Small models running locally (MLX, no
network) do the narrow parts. Between them sits a gate a machine can run β for a code change: the diff
stayed where it was asked to, tsc is clean, and every test that passed before still passes. Claude
never sees anything that did not survive it.
Illustrative shape, not a captured transcript β it is the real output format with the counts from the
React measurement below, and no client's file names. 11 of 12 with a
95 % interval of [0.615, 0.998] is the number; the line above is what it looks like while it runs.
Unit tests are workload #1 β the same shape with a mutation-kill gate instead β and the numbers for both are below, each with its interval.
sidecrew needs 24 GB of installed RAM. Below that a 7B cannot sit beside a normal working set
(measured), so sidecrew refuses to run and tells you why rather than quietly becoming a different
tool (ADR-0073). sidecrew doctor answers this before you plan anything.
And the other half of the ledger, measured 19β20 Sep 2026: the workers are free, the coordination is not. Opus planning costs 8,400β22,100 tokens per task, depending almost entirely on how many tasks are in the plan β see What this costs to run below. A page that says "the work is free" without saying that is telling you half of it.
Every rate on this page is an interval, never a point, and the reason is measured. Replaying a
frontier model's own diffs β which the gate finds no defect in β through the gate a second time, it
disagreed with itself on 2 of 19 candidates: D = 0.105, 95 % interval [0.013, 0.331]
(experiments/gate-error-rate/, against a rule frozen before the first replay). That rule's own
verdict at D > 0.10 is UNACCEPTABLE, and its consequence is this paragraph: no survival rate here
is published as a bare fraction, D travels beside each of them, and no two of our rates are compared
without asking whether their difference survives it.
What that does and does not damage. D counts the gate refusing changes that were fine.
Nothing in it is evidence of the gate admitting something bad. So:
The gate admits nothing that fails β stands. The gate refuses only things that fail β measured false, at roughly one evaluation in ten, and unproven at any tighter bound.
Two of the three known causes are now recorded on every verdict β memory pressure (ADR-0066) and a
baseline captured on a different calendar day (ADR-0069). The third is undiagnosed, and both measured
disagreements happened on a quiet machine, so it is neither of the first two. D was measured on
workload #2a's gate; it says nothing about workload #1's mutation gate, which is a different oracle
(Β§4.4 of the rule).
Against a decision rule frozen before the tool had produced a single verdict
(experiments/go-no-go-2a/results/REPORT.md). S is survival, A is a blind reviewer's approval of
the diff. Every cell is k/n with its exact 95 % interval; D = 0.105 [0.013, 0.331] applies to
every S.
| input | S, local 7B | S, Haiku control | A, local 7B | A, control |
|---|---|---|---|---|
| fixture | 3/3 [0.292, 1.000] | 3/3 [0.292, 1.000] | 1/3 [0.008, 0.906] | 3/3 [0.292, 1.000] |
| a commercial Nest/jest codebase | 12/12 [0.735, 1.000] | 12/12 [0.735, 1.000] | 9/10 [0.555, 0.997] | 10/10 [0.692, 1.000] |
| a commercial React/jest codebase | 11/12 [0.615, 0.998] | 12/12 [0.735, 1.000] | 9/10 [0.555, 0.997] | 10/10 [0.692, 1.000] |
Read the intervals, not the fractions: at these sample sizes 12/12 and 11/12 are not
distinguishable, and neither is the 7B from the control on A. Both real projects passed the frozen
rule by a margin of exactly zero, on a sample of ten. The survival rate decided nothing β five of
six measurable cells sit at or above 0.917, so the rule's ratio clauses were inert, and the only thing
that separated a local 7B from a network model was the approval rate.
What that approval rate is about, measured: the local model makes unrequested cosmetic edits β a
deleted docblock, a reworded comment, a stray blank line β in 4 of 23 [0.050, 0.388] sampled
survivors, against the control's 0 of 23 [0.000, 0.148]. confined β§ compiles β§ tests pass
cannot see any of it, which is the point.
Measured 18β19 Sep 2026 against the thing a colleague actually does today β asking a frontier model directly, one task at a time, with no plan and no gate. Same 19 hand-written tasks, same two codebases, four arms.
| Opus alone | Sonnet alone | sidecrew | |
|---|---|---|---|
| tokens, 19 tasks, Nest codebase | 80,131 | 88,820 | 0 |
| tokens, 19 tasks, React codebase | 95,335 | 88,680 | 0 |
| output vs Opus, Nest codebase | β | identical 19/19 | identical on 18 of 19 |
| changes delivered and verified, Nest | 19 (unverified) | 19 (unverified) | 19/19 [0.824, 1.000] survived |
| changes delivered and verified, React | 19 (unverified) | 19 (unverified) | 15/19 [0.544, 0.939] survived; 4 escalated |
The token columns are counts and are exact. The survival rows are rates, so they carry their
intervals, and D = 0.105 [0.013, 0.331] applies to both of them.
A fourth arm ran the same gate over Opus's own diffs, to separate the gate's contribution from the
worker's: 19/19 [0.824, 1.000] on the Nest codebase and 18/19 [0.740, 0.999] on the React
one. Two things follow, and the second is the more useful.
The gate found nothing wrong with a frontier model's work. Its single rejection there is a
tsc fragility it rejects from both arms β so on these tasks the gate adds no safety on top of
Opus. Its entire value is that it makes the free worker usable, not that it second-guesses the
expensive one.
And it is what separates a worker defect from a project defect. Of the four React failures, three
vanish when Opus writes the diff β genuine worker defects β and one reproduces exactly, because the
7B's output for that task was byte-identical to Opus's. Worker-attributable survival is therefore
15/18 [0.586, 0.964], not 15/19. Without that control the small model would have been blamed for
a third more failures than it caused β and note that the two intervals overlap almost entirely, so the
correction changes the attribution rather than the number.
Three things that buys you, and one it costs.
escalated, each with the
exact errors, none merged. Asking Opus directly gives you a diff and the sentence "nothing else
changed" β accurate here, and unverified.tsc). It is free and unattended,
not fast.What is not claimed. The frozen rule's verdict is WITHHELD: every task in these runs is a
rename, and the rule refuses a verdict until at least half are null guards, API migrations or
dead-code removal. The measurement vindicated that clause rather than surviving it β two frontier
models and a 7B emitted identical bytes on 18 of 19 Nest tasks, so that task set cannot rank
anything. experiments/status-quo/ has the arms, the protocol frozen before they ran, and the
defects.
| OS | macOS on Apple silicon. The local worker is MLX, which is Apple's. Not a port away β a different inference stack. |
| RAM | 24 GB installed or more for the local tier and everything this page measures. The threshold is on installed RAM, never free, so a machine that can host a worker never falls back silently. |
| Node | β₯ 20. A project whose own engines demands newer is honoured β run sidecrew under the project's Node (ADR-0049 cost a night to learn). |
| Python | mlx-lm, installed by you: pip install mlx-lm. Deliberately not bundled β external capabilities are shelled out and reported by doctor, never vendored. |
| Disk | ~4 GB for the 7B (downloaded on first sidecrew serve, pinned revision), plus ~500 MB per verification sandbox while a run is in flight. |
| The project | TypeScript with a green-ish tsc and a test suite that runs. Workload #1 also does Swift/XCTest. |
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/sidecrew)<a href="https://allmcps.com/mcp/sidecrew"><img src="https://allmcps.com/api/badge/sidecrew?style=directory" alt="Sidecrew on AllMCPs" /></a>