The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Speedread listing page.

Token-efficient code search and navigation for AI coding agents. speedread is an MCP server and CLI that gives Claude Code, GitHub Copilot, Codex, Cursor, Gemini CLI and other agents the part of a codebase a question needs, within a token budget: the function around each search hit, a large file's skeleton, a symbol's callers and implementations, or only what changed since the last read. Not whole files and bare grep hits.
Read is the wrong abstraction for coding agents. Agent code reading should be adaptive, stateful, symbol-aware and token-budgeted instead of byte-oriented. Every model call re-sends the system prompt, tool definitions and conversation so far (18–21k tokens before any code, in our evals), so an agent's cost is driven more by round trips than by bytes. ripgrep returns matches and cat returns bytes, so the agent asks again: open the file, find the function, search for the next hop. speedread returns the minimum useful unit of code for the question, with enough structure that the next call often isn't needed.
| The agent needs | Built-in tools return | speedread returns |
|---|---|---|
| where something is | file names, or bare matching lines | each hit under its enclosing function or class, with its line range (search) |
| one function in a large file | the file in 2,000-line pages, or a guessed range | that symbol's full source (read path#Symbol), or a skeleton of the file |
| callers, callees, implementations | a search per hop, then more reads | the relationship in one call, up to three levels deep (trace) |
| a file again, after an edit | the file again | only what changed, labelled by symbol (read path@etag) |
One of the ten code-question tasks, replayed from its recorded eval transcripts at recorded speed; all three trials of each condition behaved identically. The built-in grep answers with a file name, so the agent has to ask again, twice. speedread's search answers with the matching lines under their enclosing declaration. This is the second-largest saving of the ten tasks; two tasks came out about 1% worse, and across all ten, input tokens fell 35%. Full interactive report: brennengreen.github.io/speedread (also self-contained in demo/index.html), generated by demo/build.py from evals/results/.
Real agents on real repositories, with the same model (claude-sonnet-5) and harness (GitHub Copilot CLI) in both arms: built-in tools vs speedread as the reader. Every trial, transcript, grader and diff is committed, including the workloads where speedread didn't help.
| Workload (real agent, same model and harness) | Trials per arm | Input tokens | Model time (median) | Quality |
|---|---|---|---|---|
| Code questions: find, read, answer | 30 | −35% (95% CI −45 to −23%) | −47% | pass^3 90% → 100% |
| Relationship questions: callers, callees, implementations | 8 | −57% (CI −74 to −20%) | −34% (not significant) | 100% → 100% |
| Bug fixes: find, edit, run the test suite (with guidance · exclusive) | 16 | −3% · −1% (not significant) | −24% · −27% (not significant) | 100% → 100%; compression never hid the bug |
| Installed but not made the reader (Q&A · bug fixes) | 10 · 16 | +31% · +46% (higher on 9 of 10 · 8 of 8 tasks) | — | used in 0 of 26 trials |
trace in one hop. On bug fixes, editing and testing dominate the turns, and read results were about 1% of input, so tokens barely moved.pass^3 is the share of tasks whose three trials all passed. Intervals are 95% bootstrap intervals on the ratio of means (evals/stats.py). Per-suite detail: Results · method: evals/README.md · every table: evals/RESULTS.md · raw trials and transcripts: evals/results/
1. Install (macOS on Apple Silicon, Rust 1.90+; a clean build took 80 s on an M4, plus downloads):
Prebuilt binaries, a one-click Claude Desktop bundle, other platforms, and why the tap name: Install.
2. Add it to your agent as the reader, not as one more tool. Installed alongside the built-in tools with no guidance, it went unused and made runs more expensive (above).
VS Code, Cursor, Codex, Gemini CLI, Zed and Claude Desktop: Configuration. Where the built-in tools can't be removed, add the reading instructions to AGENTS.md, CLAUDE.md or .github/copilot-instructions.md.
3. Or try it by hand in any repository:
Four tools over MCP, mirrored by the CLI:
| Primitive | Job | Returns |
|---|---|---|
| map | locate structure | budgeted repo tree with line counts, importance-weighted; top-level symbols on request |
| search | locate text | ripgrep's engine; every hit grouped under its enclosing function or class, with line range |
| trace | locate relationships | callers, callees, references, implementations: syntactic and receiver-aware |
| read | obtain exact evidence | batched targets and path#Symbols under one token budget; path@etag returns only what changed |
read: batched, budgeted, symbol-awareOne call takes any mix of targets. They share one token budget (default 8,000).
| Target | Returns |
|---|---|
src/app.ts | The whole file. If it doesn't fit, a skeleton: signatures, types and docs, with bodies collapsed as A-B ⋯. If that's still too big, an outline. Never a blind cut. |
src/app.ts:120-180, src/app.ts:120 | Those lines; a single line (or file:line:col from a compiler error) returns the enclosing function or class. |
src/app.ts#handleRequest, #Server.start | That symbol's full source, including docs and decorators. #Name alone finds the definition anywhere. |
README.md#Install, package.json#scripts | A Markdown section, or a JSON, YAML or TOML key. |
src/**/*.test.ts | A glob (.gitignore-aware); large sets degrade largest-first. |
src/app.ts@<etag> | Only what changed since the version whose etag appeared in a header. |
A real skeleton of flask's 1,628-line app.py (excerpt) costs 3.4k tokens, against 21k for the file:
Symbol-aware re-reads. After an edit, path@etag returns unchanged, the appended tail for a growing log, or a diff that names what changed. Hunks carry git-style function context, and mode=outline returns only the symbol summary. From tests/mcp.rs:
A signature edit reads f3 [25-27]: signature changed: `pub fn f3() -> u32` → `pub fn f3(k: u32) -> u32` ; a new function reads g [133-135]: added `pub fn g() -> u8` .
Etags are 64-bit. An etag is the full 64-bit xxh3 of the content, printed as 16 hex digits, and snapshots of what the agent has seen live in a 256 MB LRU keyed by it. Two different contents would have to collide in 64 bits to alias. Across 100,000 snapshots in one session that chance is about 3 × 10⁻¹⁰. Identical contents share an etag, which is correct. Shorter tags are rejected rather than prefix-matched.
search: hits grouped by enclosing symboloutput=symbols returns the full source of every enclosing function in the same call. output=files returns paths with counts.
trace: relationships, not textThe expensive agent loop is search → open → search again to follow a call chain. trace does it in one call. Real output on gin (abridged):
Its four directions:
callers: call sites grouped by calling function; depth 2–3 builds the tree.callees: each call in the body, resolved to its definition.refs: every use, including imports and type mentions.impls: subclasses and trait, protocol and interface implementations. Go interfaces are matched structurally, by method sets. A method target lists each override.Comments and strings are excluded by tree-sitter. Same-named definitions are told apart by receiver, enclosing class, Go package and file. Sites that stay ambiguous are marked ?, never silently merged. It is syntactic, with no type inference. See Limitations.
map: a budgeted overviewA .gitignore-aware tree with line counts. Directories expand by importance until the budget (default 3,000) is spent: source first, then hidden, test or vendored trees. symbols=true adds each file's top-level definitions. Symlinks are listed as name → target and never followed.
Anthropic's Code execution with MCP argues for filtering data before it reaches the model. The four tool definitions cost ~1.4k tokens in total. The CLI mirrors them and adds JSON Lines for agents that script:
skills/speedread/SKILL.md packages the workflow as an Agent Skill.
Every number here comes from an eval in evals/, built to Anthropic's Demystifying evals for AI agents: explicit tasks, repeated trials, deterministic outcome graders, pass@k and pass^k, balanced task sets with controls, isolated trials and transcripts read. Raw data, every transcript included, is committed. All agent trials use claude-sonnet-5 via GitHub Copilot CLI, with exact token counts from the harness's usage log. The summary table covers all four real-agent workloads, not just the best one.
Suite 3: 10 questions with version-specific answers, over 6 repositories and 5 languages, with 3 trials each. Cost fell on 10 of 10 tasks, and tool results were about the same size in both conditions: the saving is two fewer round trips per answer. As the reader, speedread was used in 27 of 30 trials; the 3 exceptions were a 41-line go.mod, which the agent read with cat. This suite ran before trace, symbol diffs and the content-aware estimator existed.
Suite 3b asks for two-hop callers, resolved callees, Go interface implementations (structural) and Rust trait implementations. Two of the four are answerable with one good grep; they are the controls. Unprompted, the agent chose trace in 7 of 8 trials. On the two-hop question, built-in tools took 7–14 tool calls and 186k–277k tokens. With speedread, the agent called trace with depth: 2, checked one more function and answered: 2 calls, 61k tokens. The sample is small (8 trials per arm), so treat the size of the effect as approximate.
Transcript review changed this suite's grader. Both baseline trials of the two-hop task excluded BasicAuth, arguing that its AbortWithStatus call sits inside the closure BasicAuthForRealm returns, which the router invokes as a value. That is a defensible reading, so the grader now accepts both answers. trace attributes calls inside closures to the enclosing named function; this is listed under limitations.
Suite 4 injects 8 real regressions into gin (Go) and flask (Python), gives the agent a symptom-only bug report and grades by the repository's full test suite, with tests unmodified. Every task is verified to fail as injected and to pass with the reference fix. The four conditions are:
That is 64 trials:
trace was never called in these 32 trials: fixing a bug from its symptom needed search and read, not a call graph.Suite 1 covers 35 reading scenarios on real repositories. Each is graded for information sufficiency: a smaller answer that drops what the task needs fails. It shows what the primitives compress, and several comparisons are structurally favorable: whole-file reads and 2,000-line paging. The fair comparison is the best-case baseline, which greps for the name and reads exactly the function: −46%, in one call instead of two. The real-agent suites above are the evidence for agent performance.
Budgets are in tokens, but no client tells a server its tokenizer. Suite 2 attacks the estimator with SVG path data, JSON, lockfiles, minified JS, real CJK docs, emoji, base64, hex dumps and numeric tables. Under o200k, cl100k and the legacy Claude tokenizer, a fixed 2.6 bytes/token put 9.5% of reads over budget, the worst at 1.81×. The content-aware estimator (src/tokens.rs) models tokens from character classes and is fitted to the stricter of o200k and legacy Claude. It puts 0 of 525 reads over budget (worst 0.99×).
Calibrated against a production tokenizer. Offline tokenizers are proxies. The harness logs exact input tokens per model call, so the real cost of each tool result can be recovered from consecutive calls (Suite 2b). On claude-sonnet-5, real counts run 1.22× the estimate at the median, and 1.36× for speedread's own output, which matches Anthropic's note that Claude 4.7+ tokenizers produce ~30% more tokens. Clients that identify as Claude therefore get a Claude profile that scales budgets 1.4×. Force it anywhere with SPEEDREAD_TOKENIZER=claude; openai and legacy are the other profiles.
Speed is not the headline; returning less is. It still matters that doing more work per call, like parsing, grouping and budgeting, doesn't cost latency. Measured on an Apple M4 (4P + 6E cores), warm cache, medians:
| Task | speedread | ripgrep default | ripgrep -j4 |
|---|---|---|---|
| Walk vscode (19,167 files) with sizes and mtimes | 26 ms | 29 ms (names only) | 30 ms |
| Search vscode for a literal | 103 ms | 294 ms | 109 ms |
Read a 742 KB, 21k-line .d.ts → skeleton (cold process) | 17 ms | ||
#createDecorator definition lookup across vscode | 332 ms |
Search runs at parity with ripgrep when ripgrep is told to use only the performance cores. The 2.9× gap to ripgrep's default comes from threads spilling onto efficiency cores on this chip: kernel time grows 6.6×. speedread sizes its pool from hw.perflevel0.logicalcpu.
Other MCP servers cover parts of this. The official filesystem server batches whole-file reads and limits them by line count. Serena is symbol-aware through language servers. ast-grep MCP does structural search, claude-context searches by embeddings, and repomix packs a whole repository into one file. speedread combines batched, symbol-aware reads under one token budget with diff-only re-reads and relationship queries, and needs no embeddings or language server. The feature table, with sources, is in docs/RESEARCH.md.
From source (Rust 1.90+, via brew install rust or rustup; the tree-sitter grammars also need a C compiler, which Xcode's Command Line Tools provide):
This installs one native binary, ~/.cargo/bin/speedread, with no runtime dependencies. --locked builds the dependency versions in Cargo.lock, which CI tests.
Prebuilt, macOS on Apple Silicon: each release attaches the binary and its SHA-256.
The binary is not notarized. curl doesn't set macOS's quarantine flag; if you download it with a browser instead, clear it with xattr -d com.apple.quarantine speedread.
Claude Desktop, one click: download speedread-aarch64-apple-darwin.mcpb and open it. Claude Desktop asks which folder speedread may read; reads outside it are refused.
Homebrew: brew install brennengreen/tap/speedread builds from source. Use the full name: plain brew install speedread installs a different program, an RSVP speed-reading tool from homebrew-core that also installs a speedread binary.
MCP Registry: listed as io.github.brennengreen/speedread (mcp-name: io.github.brennengreen/speedread), so registry-aware clients can find and install the bundle. Agents installing speedread for you can follow llms-install.md.
New versions are published as releases with notes; to be notified, use Watch → Custom → Releases.
| Platform | Status |
|---|---|
| macOS on Apple Silicon | Built, tuned and tested: CI runs the test suite on macOS 15 with both directory walkers. |
| macOS on Intel | The same code. The test suite passes as an x86_64 build under Rosetta 2; not yet tested on Intel hardware or in CI. |
| Linux | The test suite passes on Ubuntu 24.04 (x86_64), and CI runs it on every push. macOS-specific code is compiled out and the portable walker (the ignore crate) is used. Not tuned or benchmarked there, and no prebuilt binary yet: install with cargo. Reports from other distributions and arm64 are welcome. |
| Other Unix | Untested; the Linux code path applies. |
| Windows | Not supported: the code uses Unix-only APIs. WSL2 has Linux's status. |
The speed figures in Results are from an Apple M4.
Availability is not adoption. Installed next to the built-in tools with no guidance, speedread was used in 0 of 26 trials across two suites. Those runs also cost more than not installing it (+31% and +46% input tokens), because its tool definitions ride along on every model call. One sentence of guidance took adoption to 16 of 16. So configure speedread as the reader, not as one option among many.
GitHub Copilot CLI: add the server, then remove the built-in readers. Edit and bash stay.
Claude Code:
Claude Code caveat, quantified. Claude Code's
Edit/Writerequire a prior nativeReadof the file; MCP reads don't count (claude-code#32214, closed as not planned). speedread can't remove that read. In the bug-fix suite, speedread's agents edited files they had only seen through speedread. A full defaultReadof each costs 2k–21k tokens (median 9.8k) and is then re-sent on every later turn. Adding it (an upper bound) moves preferred from −3% to +10% input tokens against baseline, and exclusive from −1% to +19%. The baseline doesn't change, because it read those files natively anyway. A rangedReadof the edit site is probably enough to satisfy the check, but that's unverified. In Claude Code, expect speedread to save turns and time on edit-heavy work, not tokens; the exploration and relationship savings above still apply.
VS Code, Cursor, Codex, Gemini CLI, Zed, Claude Desktop: see configuration. Add this to AGENTS.md, CLAUDE.md or .github/copilot-instructions.md:
speedread serves MCP over stdio (speedread mcp). Workspace roots come from --root <dir> (repeatable), then $CLAUDE_PROJECT_DIR, then the client's MCP roots (VS Code/Cursor folders, Claude Code --add-dir), then the current directory. Usually no flags are needed.
| Client | Config |
|---|---|
VS Code (.vscode/mcp.json) | { "servers": { "speedread": { "type": "stdio", "command": "speedread", "args": ["mcp"] } } } |
Cursor (~/.cursor/mcp.json) | { "mcpServers": { "speedread": { "command": "speedread", "args": ["mcp"] } } } |
Codex CLI (~/.codex/config.toml) | [mcp_servers.speedread] · command = "speedread" · args = ["mcp"] |
Gemini CLI (~/.gemini/settings.json) | { "mcpServers": { "speedread": { "command": "speedread", "args": ["mcp"] } } } |
Zed (settings.json) | { "context_servers": { "speedread": { "source": "custom", "command": "speedread", "args": ["mcp"] } } } |
| Claude Desktop | the .mcpb bundle (one click), or the absolute binary path plus "args": ["mcp", "--root", "/path/to/project"] |
GUI apps may not inherit your shell's PATH; use the absolute path from which speedread.
Environment variables:
SPEEDREAD_TOKENIZER=claude|openai|legacy pins the budget calibration (default: detect from the client, else legacy).SPEEDREAD_THREADS sets the worker count (default: performance cores).SPEEDREAD_WALKER=portable uses the portable walker.trace is syntactic. It uses tree-sitter plus name resolution by receiver, class and package, with no type inference, so x.f() on an unknown receiver matches every f, marked ?. Calls inside closures and lambdas are attributed to the enclosing named function, and functions passed as values aren't calls. The next step is an optional LSP/SCIP layer behind the same trace interface, for exact references, overrides and call hierarchies.context tool that picks map, search, trace or read itself are next.Each of these is written up with its scope, the skills it needs and a suggested first step in ROADMAP.md.
Read-only by construction: there are no write tools. Paths are canonicalized and must lie inside a root. Dependency caches (~/.cargo/registry, ~/go/pkg/mod, SwiftPM checkouts, SDKs) are readable; --no-deps turns that off and --unrestricted lifts all path limits. macOS privacy protections still apply. Report vulnerabilities privately; see SECURITY.md.
getattrlistbulk(2) walker gets name, type, size, mtime and flags for a batch of entries in one syscall: 1.8× faster than readdir+lstat.QOS_CLASS_USER_INITIATED.IOPOL_TYPE_VFS_MATERIALIZE_DATALESS_FILES).memchr, simdutf8, xxh3 and mimalloc.The research behind every design choice, with sources, is in docs/RESEARCH.md.
Issues, eval results and pull requests are welcome, including results where speedread doesn't help. Setup questions and ideas go in Discussions. Useful places to start:
CONTRIBUTING.md covers setup, tests, adding a language and one rule specific to this project: tool descriptions and server instructions are read by the model, so changing them can change adoption, and they need an eval. ROADMAP.md lists scoped work, including which items suit a first contribution.
MIT © 2026 Brennen Green