Searchable agent memory: BM25 corpus recall, original-text injection, remember/forget capsules.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
π‘ Paste into ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows)
Inspect callable tools, capabilities, and parameters exposed to AI agents by Mcp Vl Msa Rs.
msa_indexIndex a document; existing chunks for `doc_id` are replaced.
msa_searchTop-k chunks, score normalized 0.0β1.0.
msa_fetch_docFull original text of a document.
msa_deleteRemove a document and all its chunks.
msa_list_collectionsCollections open in the registry.
msa_statsPer-collection statistics (exact `num_documents` / `total_tokens`).
A searchable long-term memory for AI agents, exposed as an MCP stdio server.
Index documents, notes and past conversations into collections; retrieve the
top-k relevant chunks for a query and inject the original text back to the
model; add or drop agent memories with msa_remember / msa_forget. Pure
Rust, BM25 over tantivy, zero ML
deps in the default build; optional in-process dense rerank.
Any MCP client (Claude Code, Codex, or anything speaking MCP stdio) gets the same memory: a queryable corpus that survives across sessions and model swaps, with no cloud account and no embedding service required. Use it to give an agent durable recall over a knowledge base, a docs tree, or its own chat history β retrieval that returns the original text, not just embeddings.
It is one half of a two-part memory: this server is the library (corpus recall), its companion mcp-memory-rs is the notebook (curated state). An agent that swaps models loses neither.
The name: msa is the retrieval pattern it borrows from the Memory Sparse
Attention paper (arXiv:2603.23516) β an extrinsic approximation, not the
neural model; distinct from MiniMax's MSA-architecture LLMs, which are
intrinsic (in-model) generators. vl is for Vivling (codex-vl), its first
adopter β but the server is fully AI-agnostic and depends on nothing from it.
Status: v0.4 β hybrid sparse+dense optional.
The original Memory Sparse Attention paper (EverMind-AI) describes an end-to-end trainable sparse attention layer over chunk-pooled KV caches. That is a neural artifact and is not portable to a pure-Rust MCP server. What is portable, and what this repo aims to deliver, is the MSA macro pattern:
P=64 words by default, mirroring the paper).msa_search returns chunks, msa_fetch_doc returns the full document.Design and rationale are documented in the project notes (negative results, gate methodology); see docs/NEGATIVE_RESULTS.md.
Retrieval changes are decided on pre-registered, paired deltas with
bootstrap confidence intervals β not on absolute scores. Workloads: HotpotQA
(extractive QA), MLDR-it (long-doc retrieval, Italian), LongMemEval-S (500
conversational-memory questions). Full methodology, acceptance gates and
refuted hypotheses live in docs/NEGATIVE_RESULTS.md.
Headline measurements:
dense_alpha, off by default) for re-testing as
encoders improve.msa_fetch_doc after msa_search): +14.6 F1
exactly on the stratum where snippets miss the content.Reproduce:
| Tool | Since | Description |
|---|---|---|
msa_index | v0.1 | Index a document; existing chunks for doc_id are replaced. |
msa_search | v0.1 | Top-k chunks, score normalized 0.0β1.0. |
msa_fetch_doc | v0.1 | Full original text of a document. |
msa_delete | v0.1 | Remove a document and all its chunks. |
msa_list_collections | v0.1 | Collections open in the registry. |
msa_stats | v0.1 | Per-collection statistics (exact num_documents / total_tokens). |
SearchFilter | v0.2 | Metadata filter (where_eq/where_in/created_*), post-retrieval. |
msa_search_iterative | v0.3 | Memory Interleave with server-side cursor; dedups across rounds. |
msa_drop_session | v0.3 | Force-evict a Memory Interleave session before TTL. |
dense_alpha on msa_search | v0.4 | Hybrid BM25 + cosine rerank. Requires --features embeddings + [embeddings] config. |
msa_remember / msa_forget | v0.4 | Agent-memory surface: enrich + low-signal gate + content-hash dedup; standard metadata (kind / source_id / created_at). |
msa_sync_path | v0.4 | Mirror a directory into a collection (filesystem source; blake3 delta sync). |
Prebuilt binary (recommended) β download the archive for your platform from the latest release, extract, and point your MCP client at the binary:
Prebuilt targets (Linux + Android): x86_64-unknown-linux-gnu,
x86_64-unknown-linux-musl, aarch64-unknown-linux-gnu,
aarch64-unknown-linux-musl (edge / ARM / Termux), aarch64-linux-android.
macOS: no prebuilt binary is shipped (it would need Apple code-signing).
Install from source instead β cargo install below compiles it on your Mac in
one command, no signing needed.
From source (Rust toolchain) β --locked is required (the workspace
Cargo.lock pins a working time / tantivy-common resolution; a fresh
resolve breaks the build), and mcp-msa-server is the package name (the
binary it installs is mcp-vl-msa-rs):
Add [embeddings] to MCP_MSA_CONFIG to activate dense rerank. Without
this section the server stays in BM25-only mode even when the binary was
built with --features embeddings.
The production backend is candle-modernbert: the encoder runs in-process
(Candle), offline-deterministic, from a local model bundle β no daemon, no
network at runtime, no automatic downloads. Prepare the bundle once with
scripts/prepare-granite-r2-97m.sh.
A transitional backend = "ollama" (HTTP to an Ollama-compatible service)
still exists but is deprecated and scheduled for removal in v0.6 β do not
build new setups on it.
The AI client opts into hybrid scoring per-call by passing dense_alpha
to msa_search (or any future tool that supports it). dense_alpha = 1.0
(default) is BM25-only; 0.0 is dense-only; intermediate values are a
linear blend Ξ±Β·bm25 + (1-Ξ±)Β·((cos+1)/2). Cosine is shifted to [0,1]
so it composes linearly with the already max-normalized BM25 score.
Example ~/.codex/config.toml entry:
Equivalent ~/.claude.json entry for Claude Code:
instructions
text. The tool descriptions and request-field descriptions are self-contained,
so a model can work from those alone.msa_search, msa_fetch_doc, msa_stats,
msa_list_collections, msa_manifest, msa_search_iterative,
msa_interleave_round) carry the readOnlyHint annotation, which lets a
gating client auto-approve them.default_tools_approval_mode
(above) so tool calls are not blocked on a prompt.Each collection is an independent tantivy index. Collection names are validated
(rejected if they contain path separators, .., etc.) so a collection cannot
escape the root.
Shipped:
SearchFilter (where_eq / where_in / created range), post-retrieval.msa_search_iterative Memory Interleave with server-side cursor + TTL'd MsaSession registry.embeddings, Ollama backend, per-call dense_alpha; agent-memory surface (msa_remember / msa_forget); filesystem source metadata (created_at / source / ext / dir) at index time; exact num_documents / total_tokens in msa_stats; msa-bench reproducible benchmark crate; prebuilt-binary packaging.Next (not yet built):
SearchFilter runs post-retrieval; fine for
normal corpora, but a pre-filter would help when selectivity is high on a very
large index).codex-vl) β the first downstream consumer: this server is
its long-term memory.Apache-2.0. See LICENSE.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/mcp-vl-msa-rs)<a href="https://allmcps.com/mcp/mcp-vl-msa-rs"><img src="https://allmcps.com/api/badge/mcp-vl-msa-rs?style=directory" alt="Mcp Vl Msa Rs on AllMCPs" /></a>