Retrieval over the folder the session opened: files, search, read, code_pack, write_note.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent — or use 1-click editor setup below.
One-click editor setup isn’t available for this listing yet — we don’t have a confirmed install command, and we’d rather show nothing than point your editor at the wrong package or host. Follow the project’s own setup instructions, linked above.
Lightweight, local evidence retrieval tools for code and research
Silica locates the source, symbol, page, or passage you and your agents need. Maximum signal, minimum machinery.
Silica indexes the markdown, code, PDFs and office files under one root and serves them to Claude Code, Cursor, Hermes, Codex, OpenCode and every other popular harness, as MCP tools or as shell commands that print the same JSON. A hit is a path, a section, a line, a window of text and the numbers to judge it by. The harness owns the loop.
Three commands, then a question asked the way you would ask any other. The
harness calls silica_search, and the answer carries the file, the page and
the line it came from.
The second line adds the dense leg: numpy and a static model2vec model, no
torch and no GPU. It stays inert until the model is named and the sections
are embedded — the last stanza of the Quickstart. Take it when the questions
are paraphrases that share no words with the text; an exact term or an
identifier is answered by the lexical leg either way, and only that leg
reports terms_absent.
pipx works the same. The package is silica-core, the command is
silica, the tools are silica_*.
In any folder of markdown, code, PDFs or office files:
Nine arXiv papers, indexed in 2.1 s. The hit names the file, the page and the passage; the page beside it is the check.
| Tool | Shell | Returns |
|---|---|---|
silica_files | silica files | the inventory and what the index did with each file: indexed, changed, excluded, failed, unconverted |
silica_search | silica search | ranked passages: path, section, line, BM25, matched terms, coverage, and the query terms absent from the corpus |
silica_read | silica read | a slice by lines or by heading (a page, in a PDF), the outline, and a version to carry forward |
silica_code_pack | silica code-pack | an AST context pack for one source file inside a character budget |
silica_write_note | silica write-note | one atomic write, linted for structure and unresolved wikilinks |
In a source tree every function, method, class and constant is its own unit:
a hit's section is the symbol, span its lines, and
silica_read(path, section=…) serves the body. For a symbol whose name is
known, grep wins; for a question that names none, the search comes first:
the tool description and the block silica setup claude writes say so. The
contract, the reply shapes and the acceptance checks are in
TOOLS.md. Nothing needs an API key or a network.
A ranked list always has a top, even when the corpus does not answer. Three fields say how much the result is worth:
coverage: the share of the query's idf mass the hit's matched terms carry. Near 1, every rare term matched; near 0, only common words did.terms_absent: query terms that occur nowhere in the corpus.matched_terms: the words this hit actually contains.
raft consensus log replication over the same nine papers: raft and
consensus occur in none of them, coverage falls to 0.19, and the top hit
is about data replication. On 254 papers the top hit of an answered question
carries 0.69 to 1.00; a question the corpus does not cover, 0.44. Silica
exposes the signals; the harness decides whether to stop, read or rephrase.
nDCG@10 on BEIR SciFact · NFCorpus, the same documents and queries for every arm. BEIR's published BM25 baselines are 0.665 · 0.325. Silica's lexical index needs no model; the others serve lexical search from an index that also holds embeddings.
| Mode | Silica | zvec-grep 0.2.2 | ck 0.7.11 |
|---|---|---|---|
| Lexical | 0.662 · 0.311 | 0.649 · 0.297 | 0.630 · 0.289 |
Hybrid, same potion-retrieval-32M embedder | 0.675 · 0.328 | 0.672 · 0.330 |
Code, on the twenty SWE-QA questions zvec-grep publishes for its own benchmark, same embedder, k = 10, scored on the files and symbols the reference answer rests on.
| Arm | file hit@5 · @10 | file MRR | symbol hit@10 | symbol recall | chars returned |
|---|---|---|---|---|---|
| Silica, hybrid | 0.85 · 0.90 | 0.68 | 0.75 | 0.24 | 8,266 |
| Silica, vectors | 0.80 · 0.85 | 0.67 | 0.65 | 0.21 | 6,225 |
| Silica, lexical | 0.65 · 0.75 | 0.47 | 0.45 | 0.14 | 8,194 |
| zvec-grep 0.2.2, hybrid | 0.65 · 0.75 | 0.54 | 0.55 | 0.17 | 7,326 |
| zvec-grep 0.2.2, vector | 0.60 · 0.80 | 0.61 | 0.55 | 0.18 | 6,821 |
| zvec-grep 0.2.2, FTS | 0.45 · 0.60 | 0.36 | 0.45 | 0.11 | 6,690 |
On BEIR the two hybrids tie at the 95% interval: the same vectors rank the same, with no daemon and no vector store. On code, Silica's hybrid file MRR is +0.135 over zvec-grep's hybrid (95% interval +0.01 to +0.27), paired per question. The fusion also gains +0.21 MRR and +0.30 symbol hit over Silica's lexical arm; no reranker or graph expansion is involved.
On this measured scope, Silica is a compact, local, SOTA-competitive retriever: it matches zvec-grep on BEIR and leads the paired SWE-QA code-localization replay with the same embedder.
Retrieval matters only if the agent does less work without losing the answer. These are separate experiments and are not pooled:
| Workload and arm | Runs | Quality | Search used | Turns | Tool calls | Seconds | Warm cost |
|---|---|---|---|---|---|---|---|
| Repository, search-first contract | 20 | Judge 59.7 | 17/20 | 4.7 | — | 24 | $0.197 |
| Repository, same plugin without contract | 20 | Judge 50.6 | 0/20 | 6.5 | — | 28 | $0.180 |
| Documents, resident Silica tools | 12 tasks | 12/12 correct | 12/12 | 3.9 | 2.9 | — | $0.20 |
| Documents, no plugin | 12 tasks | 12/12 correct | — | 5.0 | 4.0 | — | $0.22 |
The repository result is one repetition: turns improve by 1.75 (95% interval 0.55 to 3.05 fewer), while Judge and cost remain inconclusive. The document rows belong to a 144-run study over twelve questions and a 5.5M-token corpus. They establish less work on that workload, not a universal agent claim.
Corpora, intervals, per-task exceptions and reproduction commands are in benchmarks.
silica setup <client> writes the registration into the client's own config
and backs up what was there; for claude it also puts a guidance block, when
to search before grep, into ~/.claude/CLAUDE.md. silica setup --list
names the clients: claude, codex, cursor, windsurf, zed, cline,
roo, continue, goose, opencode, openhands, gemini, dsh,
hermes, openclaw, agent-zero, claude-desktop, lmstudio,
anythingllm and librechat; shell, python and generic print recipes
for anything else. The server serves the folder the client opens in;
--vault DIR or SILICA_VAULT fixes the root.
No reviews yet — be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/silica-core)<a href="https://allmcps.com/mcp/silica-core"><img src="https://allmcps.com/api/badge/silica-core?style=directory" alt="Silica Core on AllMCPs" /></a>