Model routing for AI agents: delegate bulk reads & boilerplate to cheap worker models.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
One-click editor setup isnβt available for this listing yet β we donβt have a confirmed install command, and weβd rather show nothing than point your editor at the wrong package or host. Follow the projectβs own setup instructions, linked above.
A decoupled, zero-dependency, universal implementation of the Shunt model-routing pattern (originally conceived by Spotify Engineering).
Model-Shunt allows AI coding agents (Antigravity, Cursor, Windsurf, Claude Code, Aider, OpenHands, etc.) to delegate token-heavy I/O (bulk file reading/code analysis) and repetitive boilerplate generation (tests, mocks, stubs, configs) to fast, economical, or local worker models (Gemini 2.5 Flash, Groq/Llama, Ollama, DeepSeek, GPT-4o-mini). This cuts primary agent token consumption by up to 90% while keeping the main context window clean.
urllib, json, re, argparse). No pip install, no virtual environment, and no npm required.gemini-2.5-flash, llama-3.3-70b-versatile, gpt-4o-mini).qwen2.5-coder:latest, gemini-2.5-flash, deepseek-chat).ARG_MAX Limits: Unlike naive implementations that pass file contents as CLI arguments (capped at ~128 KB on Linux), Model-Shunt streams corpus data over stdin, allowing analysis of hundreds of thousands of lines without buffer overflows.N|): Automatically prefixes every line in file blocks with its 1-based index, forcing worker models to cite verifiable, exact line numbers instead of hallucinating locations.bulk_read payload exceeds the direct limit (SHUNT_MAX_DIRECT_TOKENS, default ~200k tokens), Model-Shunt automatically splits the corpus into chunks, maps the question over each chunk (preserving absolute N| line numbers), and reduces the extracts into one cited answer. Giant single-line files (minified JSON/JS) are sliced by characters with explicit position markers. Rate-limit pacing waits out provider quota windows instead of failing.Configure your worker model via environment variables or a config.json file (placed in ~/.config/model-shunt/config.json or in the project root):
config.jsonTip: Setting
"model": "auto"(or passing--auto-modelin the CLI) will automatically inspect the provider's active models and pick the optimal one for reading vs writing.
Security: Do not put your API key in
config.jsonβ use environment variables instead (e.g.GEMINI_API_KEY,GROQ_API_KEY, orSHUNT_API_KEY). Anapi_keyfield exists as a last-resort fallback, but keeping secrets out of files is strongly recommended.
Model-Shunt provides a standard stdio MCP server exposing three tools:
get_available_models(provider?): Queries the provider endpoint and recommends reader and writer models. If discovery fails or is unsupported, returns built-in recommendations whose current availability is not verified.bulk_read(question, file_paths, model?, provider?): Reads large or multiple files and outputs concise, structured bullets with exact line citations. A citation name (N|k) is dropped when source line k does not contain that name.code_write(spec, reference_path, target_path?, model?, provider?): Generates code using a specification and reference file. Returns the code when target_path is omitted; otherwise creates missing parent directories and writes the code, overwriting an existing target file.bulk_read sends the selected files and question to the configured worker provider; code_write sends the specification and reference file. The provider can be remote or local (such as Ollama). Token savings depend on the input and response.
The MCP server runs locally over stdio. Set provider credentials through environment variables. File access follows resolved paths, so symlinks pointing outside the allowed workspace roots are rejected. Invalid requests return structured errors and leave the server available for subsequent calls.
MCP Registry name: mcp-name: io.github.yasmanycastillo/model-shunt
Universal one-liner (detects uv / pip / pipx / npm, installs the model-shunt command, and registers it with Claude Code if present):
Manual alternatives:
Any MCP client (Cursor, Windsurf, Antigravity, Claude Desktop, etc.) β add to its MCP settings. No clone, no absolute paths:
Fallback (offline / no uv / no npx): run straight from a clone with Python 3.9+ β replace
"command"/"args"with"command": "python3", "args": ["/absolute/path/to/model-shunt/src/model_shunt/server.py"].
Security: by default
bulk_read/code_writeonly operate on files inside the server's working directory (the agent workspace). SetSHUNT_ALLOWED_ROOTS(PATH-style list) to expand the sandbox.
| Variable | Default | Purpose |
|---|---|---|
SHUNT_MAX_DIRECT_TOKENS | 200000 | Payloads above this estimated size switch to map-reduce |
SHUNT_CHUNK_CHARS | 600000 | Chunk size in characters (~150k tokens) |
SHUNT_CHUNK_RETRIES | 3 | Retries per chunk on rate limits |
SHUNT_CHUNK_RETRY_DELAY | 60 | Seconds to wait out a provider quota window (free-tier TPM) |
For agents supporting pre-execution hooks (e.g., Claude Code, custom agent loops):
check-file-size):
SHUNT_MIN_LINES), the hook blocks the call and instructs the agent to delegate to bulk-read.offset and limit are allowed, preserving surgical context for code editing.check-bash-read):
cat, less, or more on large files directly in the terminal context.You can also use Model-Shunt directly from the command line or from agent bash sessions:
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/model-shunt-2)<a href="https://allmcps.com/mcp/model-shunt-2"><img src="https://allmcps.com/api/badge/model-shunt-2?style=directory" alt="Model Shunt on AllMCPs" /></a>