The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Mnemostack listing page.
Self-hosted hybrid memory & retrieval for AI apps.
mnemostack is a durable retrieval layer over your own Qdrant (and optional Memgraph): semantic, keyword (BM25), temporal, and graph recall, fused with Reciprocal Rank Fusion and refined by an 8-stage ranking pipeline — with payload filters for multi-tenant isolation, optional LLM answer synthesis (confidence + citations), and an ingest path that enriches and projects structured fields. One recall(query) call, usable as a Python library, an HTTP service, or an MCP server.
Flagship use case — durable memory for AI agents. Long-running agents hit the same wall: context gets compacted, sessions restart, useful decisions disappear, and the next run pays the re-orientation tax again. mnemostack gives them a persistent memory layer to query when the context window is not enough — durable, searchable, scoped, and explainable, not just embedded and hoped for.
The same engine backs other retrieval-heavy work: RAG over mixed corpora, multi-tenant or per-user knowledge stores, and time-aware search backends — anywhere pure vector similarity falls short on its own.
Status: Actively developed — public API is stable; new functionality lands additively in minor releases. Breaking changes are rare and called out in CHANGELOG.md.
The fastest on-ramp is the MCP server — it gives Claude Desktop, Claude Code, Cursor, ChatGPT, or another MCP-capable agent durable memory in a few commands. Building an app instead of wiring up an agent? Use the HTTP API or the Python library over the same collection.
Run a local Qdrant for the vector store:
Optional: run Memgraph for graph-backed memory:
Claude Desktop config example:
Claude will then be able to call mnemostack_search, mnemostack_answer, and graph tools.
Index a folder of notes, docs, transcripts, or project context:
--recreate drops the existing collection, so it asks for confirmation first; pass --yes to skip the prompt (required in scripts/CI — non-interactive runs without it exit with code 2).
For a running app or assistant, use the streaming Ingestor API shown below to store messages as they arrive.
From an agent, ask a memory-style question and let the MCP tools retrieve the right facts.
From the shell, test the same collection directly:
Vector search answers "what sounds similar?" Real retrieval over a growing corpus needs to answer "what actually matters for this query?" — which takes exact matches, semantic similarity, relationship tracking, recency, user/project scope, and feedback from past recalls. Mnemostack uses hybrid retrieval so recall is reliable instead of embedding roulette. (Agent memory is the most demanding version of this problem, which is why it's the flagship use case.)
Agent and chatbot memory (flagship):
filters so one user never sees another's history.Beyond agents — the same engine as a retrieval backend:
answer() (with confidence and source citations) over it from the CLI, HTTP, or library.filters isolate each tenant's data inside every retriever, or enable service-key auth (serve --auth) for a hard, key-resolved tenant boundary with optional per-tenant quotas (see the HTTP API).mnemostack mcp-serve, connect your agent, and use memory tools from the chat/runtime you already use.mnemostack serve and call /recall, /answer, /feedback, the memory write/lifecycle endpoints (/memories, /invalidate, /triples), /health, /metrics, or /docs from any language.Think of it as a storage hierarchy for agent memory:
recall(query) = page fault handler. When the agent needs something that isn't in the current context, it pulls the exact fact from storage with a single hybrid query — not a grep, not a reload of the whole corpus.The practical effect: you stop re-explaining your project to the agent after every /compact. You stop losing momentum to the re-orientation tax that shows up in any agent with session compaction. mnemostack solves it at the library level, not tied to any single agent runtime.
On each recall(query): the configured retrievers (Vector and Temporal by default, with BM25 and Memgraph when configured) run in parallel and return ranked lists. Reciprocal Rank Fusion merges them. The optional 8-stage pipeline can reweight results using query classification, exact-token rescue, gravity/hub dampening, freshness, inhibition-of-return, curiosity boosts, Q-learning weights supplied through its state store, and graph resurrection. An optional LLM reranker does a final ordering pass. You get a list of RecallResult with source, score, and provenance — ready to hand to a model. The list order is authoritative: score has no single scale — many stages and fallback paths write it, and a rerank changes the order without rewriting the numbers — so re-sorting by them undoes it. It is not a similarity, not a confidence, and not comparable across queries; see what score is not.

Most memory tools in the agent ecosystem pick one axis and optimize for it: simple vector similarity for RAG, framework-bound memory tied to a specific agent library, platform-level runtimes with audit and compliance features, or CLI wrappers over a single vendor's session store. Each makes sense for its scope.
mnemostack takes a different slice: it is a recall quality layer, offered as a plain Python package. Four retrievers (Vector + BM25 + Memgraph + Temporal), RRF fusion, an 8-stage pipeline, and an optional LLM reranker — composed to handle mixed workloads on the same corpus: exact-token lookups, semantic queries, temporal questions, and multi-hop reasoning, without forcing you to choose one mode over another.
We are not a replacement for your agent framework and not a full platform runtime. We are the piece that actually finds the right fact in a growing corpus. Drop mnemostack into your own Python agent or application, or let a higher-level service call recall() over a plain function boundary. The retrievers, pipeline, and reranker are individually composable — take only the parts you need.
See ARCHITECTURE.md for detailed design: pipeline stages, Qdrant schema, Memgraph temporal model, consolidation runtime, MCP tools.
/feedback or mnemostack feedback updates usefulness signals without silently training on every response.The 8-stage pipeline can use a small state store between calls (Q-learning weights, inhibition-of-return history, per-document gravity/hub counters). FileStateStore(path) persists it to a JSON file. HTTP recall applies existing state and can record inhibition-of-return exposure with --auto-record-ior; Q-learning updates only through explicit /feedback calls. CLI/MCP recall still apply existing state but do not collect feedback automatically. For deterministic benchmarks, call build_full_pipeline(enable_stateful_stages=False) so IoR/Q-learning/curiosity state cannot affect scores. For multi-process servers, implement your own StateStore (three methods: get(), set(), update()) backed by Redis or your database.
Any retriever can fail (Memgraph down, Qdrant unreachable, BM25 corpus empty). Recaller logs and continues with the remaining sources. The LLM reranker is wrapped in try/except by convention — if the LLM is rate-limited, the pre-rerank order is returned. This is deliberate: a memory stack that goes dark because one component hiccuped is worse than a slightly degraded one.
One exception: query expansion runs before retrieval, so a misconfigured expansion step (query_expansion=True without an expansion_llm, or a provider error inside it) surfaces as an error instead of degrading silently — see ARCHITECTURE.md for the full fail-open contract. Degradations themselves are visible, not silent: every HTTP/MCP response carries degraded tags, and the full per-retriever trace is available opt-in via include_trace.
On LoCoMo, Mnemostack reaches 82.9% strict accuracy in our evaluation setup. The table below includes our baseline runs and externally reported numbers for context. Results depend on dataset version, configuration, judge model, scoring rules, and query type. Treat externally reported numbers as directional unless they were run with the same harness and settings.
Full LoCoMo runs use the official SNAP-Research dataset (10 samples / 1986 QA) from a clean state. Across the tables below: Strict = exact match, Combined = strict + partial. Counts in cells are correct / total.
Some LoCoMo cat_5 questions have empty ground-truth answers. Under the current scorer, these are counted as correct because there is no expected answer to match. To avoid overstating recall quality, we also report signal-only scores with those questions removed. Signal-only scores are computed on the 1,540 questions with non-empty ground-truth answers.
gemini-3-flash-preview)| Run | Strict (full) | Combined (full) | Strict (signal-only) | Combined (signal-only) |
|---|---|---|---|---|
| Baseline v0.3.0 (Vector + BM25 + 8-stage pipeline) | 76.7% (1524 / 1986) | 88.1% (1750 / 1986) | 70.0% (1078 / 1540) | 84.7% (1304 / 1540) |
Retrieval improvements (window_size=3, query expansion, top-K 25) | 82.5% (1639 / 1986) | 92.2% (1832 / 1986) | 77.5% (1193 / 1540) | 90.0% (1386 / 1540) |
| v0.4.5 + photo captions (same config as above) | 82.9% (1647 / 1986) | 92.7% (1842 / 1986) | 78.0% (1201 / 1540) | 90.6% (1396 / 1540) |
Honest numbers disclaimer.
(full)is the headline aggregate across all 1986 questions, the format vendors typically report — some publish only their strongest sub-category, we publish the full aggregate because it's what actually predicts behavior on mixed workloads.(signal-only)strips thecat_5auto-pass artifact described above, so what you read there is the real recall quality on questions that have a ground-truth answer.
Per-category breakdown (v0.4.5 + photo captions run):
| Category | Strict | Combined |
|---|---|---|
cat_1 single-hop lists | 51.4% | 88.3% |
cat_2 temporal | 79.8% | 85.7% |
cat_3 open-domain reasoning | 62.5% | 79.2% |
cat_4 multi-hop reasoning | 88.0% | 94.6% |
cat_5 adversarial open-domain | 100.0% | 100.0% |
Notes:
gemini-3-flash-preview is more accurate than the previous Gemini Flash judge on synonyms, partial matches, and empty ground truth.cat_5 questions have empty ground truth in this new run and are auto-scored as correct by the benchmark harness. That makes the new cat_5 strict score (446 / 446, 100.0%) useful for aggregate harness accounting, but not directly comparable to the historical cat_5 strict score (89.7%) from the older adversarial-question evaluation.rerank_mode does not affect these numbers.blip_caption) that LoCoMo attaches to image-sharing turns — 697 of the 1540 signal questions cite image turns as evidence, and earlier runs silently dropped that content. Answer prompts also show the time of day of each memory since v0.4.5.gemini-2.5-flash judge)| Metric | First full run | mnemostack 0.2.1 |
|---|---|---|
| Strict | 66.4% (1319 / 1986) | 67.8% (1346 / 1986) |
| Partial | 12.8% (254 / 1986) | 12.6% (250 / 1986) |
| Wrong | 20.8% (413 / 1986) | 19.6% (390 / 1986) |
| Combined | 79.2% (1573 / 1986) | 80.4% (1596 / 1986) |
By question category (combined not tracked for the first full run):
| Category | First run Strict | 0.2.1 Strict | 0.2.1 Combined | Δ Strict |
|---|---|---|---|---|
cat_1 single-hop lists | 34.8% | 34.4% | 74.1% | −0.4pp |
cat_2 temporal | 64.5% | 69.8% | 77.9% | +5.3pp |
cat_3 open-domain reasoning | 31.2% | 41.7% | 49.0% | +10.5pp |
cat_4 multi-hop reasoning | 69.2% | 69.6% | 82.0% | +0.4pp |
cat_5 adversarial open-domain | 90.1% | 89.7% | 89.7% | −0.4pp |
Last historical run: 2026-04-27, mnemostack 0.2.1, same dataset, judged by gemini-2.5-flash.
Caveat: different judges, evaluation protocols, and in some cases category cherry-picking. Vendor numbers below are taken at face value from their published material.
| System | LoCoMo correct |
|---|---|
| Hindsight (reported range) | 78–85% |
| Memobase (temporal subset) | 85% |
| mnemostack | 82.9% |
| Letta filesystem agent | 74% |
| Mem0 graph variant | ~68.5% |
| Zep (independently replicated) | 58.4% |
LoCoMo measures generic long-term dialogue recall. We also run a private needle-in-haystack benchmark on the production workload that drove the original design — a ~17k-point memory stack indexed from a long-running assistant. Queries mix exact tokens (IP addresses, tickers), telegram IDs, paraphrased facts, and temporal probes.
| Metric | Value |
|---|---|
| recall@1 | 90% (9/10) |
| recall@5 | 100% (10/10) |
| recall@10 | 100% (10/10) |
| Query latency p50 | 1.26 s |
| Query latency max | 1.70 s |
Honest numbers disclaimer. Reporting only
recall@5 = 100%would look impressive, but it would also hide the harder top-1 behavior.recall@1 = 90%is what an agent reading only the top hit actually experiences, and the gap between@1and@5is where reranker quality (or the lack of it) shows up. We publish all three so you can read the metric that matches your downstream usage.
Useful because LoCoMo's failure modes (list exhaustion, open-domain reasoning) are orthogonal to what production memory stacks actually spend time on (find the specific fact the user mentioned weeks ago). This benchmark is not in the public repo; its methodology is in benchmarks/synthetic_longhorizon.py, which is the closest reproducible approximation.
Details, category definitions, and notes on the judge protocol: benchmarks/README.md.
Build it in if you need:
Not the best fit if you only need a single call to text-embedding-3-small + cosine similarity — something simpler will do. mnemostack earns its complexity on mixed, long-horizon workloads.
Retriever abstraction — add your own sources.reciprocal_rank_fusion(weights=[...]) lets you lift sources you trust more; Recaller(adaptive_weights=True) picks a per-query-shape profile (exact-token / person / temporal / general). See the honest write-up below for where this helps and where it doesn't.search(). Not included in the default Recaller./feedback, and recall exposure logging is off unless --auto-record-ior is enabled.ScoringReranker. See docs/recipes.md for a runnable bge-reranker-v2-m3 example.BM25Retriever(tokenizer=...) for stemming / lemmatization / language routing. Core stays dependency-free; docs/recipes.md has per-language recipes.Recaller.recall_async, recall_flow_async, Ingestor.ingest_async / ingest_one_async, AnswerGenerator.generate_async, synthesize_async, plus AsyncVectorStore over the native async Qdrant client. Retrievers dispatch in parallel; five concurrent HTTP recalls finish in roughly one single-recall wall-clock.store.invalidate(ids, valid_until=...) sets bi-temporal payload keys (invalidated_at system-time, valid_until/valid_from world-time) via a cheap merge write. Recall hides invalidated facts by default; include_invalidated=True shows them and as_of="<iso>" reconstructs what was valid at a past instant from valid_from/valid_until — each optional, and invalidated_at is not read there at all, so a point-in-time view can be a superset of the default one (contract). The vector-side twin of the graph's valid_until model. CLI mnemostack invalidate <id>..., MCP mnemostack_invalidate, and — since 2.2 — HTTP POST /invalidate (with DELETE /memories for irreversible erasure), selecting either an id list or a whole source.telegram_id, handle, and precomputed name_lower so non-ASCII names match correctly (Memgraph's toLower() lower-cases ASCII only).Ingestor API — batched, idempotent, LRU-cached ingest from any Python code. Lazy iterator means large corpora ingest with bounded memory. Same (source, offset, text) → same deterministic UUID-shaped content id, so re-runs are no-ops.mnemostack index-markdown <dir> indexes a folder of markdown with structure: YAML frontmatter → payload filters, header-aware chunking with heading paths, and [[wikilinks]] / [text](https://github.com/udjin-labs/mnemostack/blob/HEAD/note.md) → File -[LINKS_TO]-> File graph edges (with a Memgraph URI). Generic for any markdown folder; Obsidian vaults work as a side effect. Depends only on the already-present pyyaml.pip install 'mnemostack[server]' gives you /recall, /answer, the write/lifecycle surface (POST/GET/DELETE /memories, /invalidate, /triples), /health, /docs, plus /metrics in Prometheus text format. See the HTTP server section below.valid_from/valid_until, query point-in-time state; graph resurrection stage recovers evicted-but-relevant memories.cat_3 inference retry with query decomposition are on by default.synthesize(entity) rolls up everything memory knows about a person, project, or topic into a structured profile (SynthesisFact / SynthesisResult, markdown or JSON). CLI: mnemostack synthesize <entity>. Optional related-entities expansion via graph and LLM summarization pass.search --tier {1,2,3} and answer --tier {1,2,3} bound output size (~50 / ~200 / ~500 tokens) so agents can pay only for the detail they actually need. Omit --tier for unchanged full output.MessagePairChunker for chat transcripts (keeps user↔assistant pairs together). The new vector.window_size config carries adjacent-turn context inside each chunk; window_size=3 was worth +5.8pp strict / +4.1pp combined on LoCoMo (v0.4.0).Recaller(expansion_llm=...) widens recall with reformulated queries; AnswerGenerator(retry_with_expansion=True) retries low-confidence answers with the expanded query and a HyDE-style hypothetical before giving up. Opt-in via --query-expansion on mnemostack answer.filters={"tenant": "a"} applies inside every retriever (exact match + ranges) on HTTP/MCP/CLI/library; results never include points outside the scope, verified by adversarial isolation tests. Filters are caller-supplied, so for a real trust boundary run the server with service-key auth (serve --auth / mcp-serve --auth): the tenant is resolved from the key (a client can't assert another's), enforced across the vector store, the knowledge graph, and per-tenant learning state. Optional per-tenant storage quotas apply at ingest, and request-rate quotas on the authenticated HTTP surface (serve --auth). Off by default. See the HTTP API section.Ingestor(enrich=callable) extracts structured facts into payloads at ingest (fail-open, --refresh-payloads updates existing collections without re-embedding); context_fields=[...] shows them to the answer LLM; rewrite_followup() resolves conversational follow-ups before recall.think is off by default (reasoning models otherwise burn the whole token budget on thoughts and return empty text); options={...} passes any generation option through.Some of the newer knobs help in specific workloads and do nothing (or mildly hurt) in others. Measured, not promised — both are opt-in by design, and the default Recaller stays classical equal-weight RRF over Vector + BM25 (+ Memgraph + Temporal when supplied).
Recaller(adaptive_weights=True) — picks a weight profile per query shape:
| Query shape | Detection | Profile (bm25 / memgraph / vector / temporal) |
|---|---|---|
exact_token | IPv4 / port / version / UUID / API-style tokens | 1.4 / 1.4 / 1.0 / 0.9 |
person | "who is", @handle, username, contact, etc. | 1.0 / 1.5 / 1.0 / 0.9 |
temporal | "when", "yesterday", "today", dates | 1.0 / 1.0 / 1.0 / 1.4 |
general | everything else | classical equal-weight RRF |
Measured on a real production corpus with 10 needle probes: recall@1 went 50% → 60%, recall@5 stayed at 90% (zero regression). On LoCoMo (pure dialogue questions, all classified general), adaptive weights had no effect — the profile simply isn't triggered. Rule of thumb: turn it on for production ops-style workloads (IPs, tickers, IDs, named entities); leave it off, or don't expect a lift, for dialogue benchmarks. Static retriever_weights={...} always wins over adaptive when both are set.
HyDERetriever — generates a short hypothetical answer via your LLM and embeds that instead of the raw query, then fuses alongside the other retrievers. Useful when the question and the stored answer use very different vocabulary (documentation corpora, FAQ-style content). On our LoCoMo cat_3 smoke (conv-43, 14 open-domain reasoning questions) it moved accuracy from 14.3% to 21.4% (+1 correct answer); on dialogue-backed memory overall it's roughly a wash. It always costs one extra LLM call per search(), so budget accordingly and treat it as a tool for specific workloads rather than a default.
Agent runtimes often wrap transcript messages in metadata envelopes before the real body, which can dominate embeddings and make unrelated turns look similar. Clean messages before chunking/indexing with strip_metadata_blocks():
Built-in profiles cover OpenClaw webchat and Telegram envelopes; pass profiles= or extra_patterns= to tune the cleanup for your runtime.
| Variable | Purpose | Required for |
|---|---|---|
GEMINI_API_KEY | Google Generative AI key | Gemini embedding + Gemini Flash LLM |
OLLAMA_HOST | Ollama server URL (default http://localhost:11434) | Ollama embeddings / LLM |
MNEMOSTACK_COLLECTION | Qdrant collection name (default mnemostack) | CLI convenience |
MNEMOSTACK_QDRANT_URL | Qdrant URL (default http://localhost:6333) | Remote Qdrant |
MNEMOSTACK_GRAPH_URI / MNEMOSTACK_MEMGRAPH_URI | Memgraph bolt URI | Graph retriever / GraphStore |
MNEMOSTACK_LLM_HOST / MNEMOSTACK_LLM_TIMEOUT | LLM endpoint (ollama: default inherits the embedding --ollama-host; openai: required base URL) and LLM request timeout | Answer / reranker / expansion LLM |
MNEMOSTACK_LLM_API_KEY | Bearer token for the openai LLM provider; unset or none = no auth header (keyless vLLM / llama.cpp) | Answer / reranker / expansion LLM |
MNEMOSTACK_PROVIDER / MNEMOSTACK_EMBEDDING_PROVIDER | Embedding provider | CLI / HTTP / MCP |
MNEMOSTACK_LLM / MNEMOSTACK_LLM_PROVIDER | LLM provider | Answer generation / reranking |
MNEMOSTACK_BM25_PATHS | BM25 corpus paths separated by os.pathsep (: on Unix) | CLI / HTTP / MCP BM25 retriever |
MNEMOSTACK_AUTO_RECORD_IOR | true/false toggle for HTTP recall exposure logging | HTTP stateful pipeline |
MNEMOSTACK_EMBEDDING_MODEL / MNEMOSTACK_LLM_MODEL | Override the embedding / LLM model name | CLI / HTTP / MCP |
MNEMOSTACK_VECTOR_HOST / MNEMOSTACK_VECTOR_COLLECTION | Aliases for the Qdrant URL / collection | CLI / HTTP / MCP |
MNEMOSTACK_VECTOR_FLOOR | Keep top-N raw vector hits in results even when fusion/rerank would drop them (0 = off) | Recall tuning |
MNEMOSTACK_RERANK_MODE | LLM reranker mode: relevant_only (default) or full_reorder | HTTP / MCP runtime reranker |
MNEMOSTACK_TOKEN_BUDGET | Default recall token budget — cut results to the ranked prefix that fits (unset = off) | CLI / HTTP / MCP recall surfaces |
MNEMOSTACK_GRAPH_TIMEOUT / MNEMOSTACK_GRAPH_HEALTH_TIMEOUT | Memgraph query / health-check timeouts in seconds | Graph retriever |
MNEMOSTACK_CONFIG | Path to the YAML config file | All entry points |
Only the providers you actually use need their keys. HuggingFace local-GPU embeddings need no keys at all. mnemostack init writes the same settings as YAML; explicit CLI flags override config/env defaults.
Fastest way to kick the tyres. No Python install, no manual Qdrant / Memgraph setup.
The mnemostack container runs the HTTP API on port 8000 by default. Interactive docs are at http://localhost:8000/docs. Use docker compose exec mnemostack mnemostack <cmd> for CLI-style operations (index, search, health) against the same stack.
Tear down with docker compose -f examples/docker-compose.yml down -v (the -v wipes Qdrant + Memgraph state).
Prefer Ollama (no cloud key needed)? Run Ollama on the host and pass --provider ollama everywhere instead of gemini. The endpoint resolves as: --ollama-host flag > MNEMOSTACK_OLLAMA_HOST env / embedding.ollama_host config > the native OLLAMA_HOST variable > http://localhost:11434 — so a client running in a container or VM can reach a remote Ollama daemon directly. An ollama LLM follows the same chain and inherits the embedding host by default; set llm.host / MNEMOSTACK_LLM_HOST only when generation lives on a different box:
Embedding uses the batch POST /api/embed endpoint (one request per batch; servers too old for it are detected once and served per-item with a loud warning). The embedding timeout (--embedding-timeout / MNEMOSTACK_EMBEDDING_TIMEOUT, default 180s) is independent of the short Qdrant liveness timeout — cold loads of larger local models are legitimately slow. Vector dimensions come from the model tables (quantization-suffix aware) or, for unknown models, a one-shot probe of the live model — there is no blind fallback dimension, so a wrong-size collection can't be created.
Behind an OpenAI-compatible endpoint (LiteLLM proxy, vLLM, llama.cpp server, an API gateway)? Use the openai LLM provider — it speaks POST {base}/v1/chat/completions, which all of them accept:
with llm.host: http://gateway:4000 in the config (or MNEMOSTACK_LLM_HOST). Both the base URL and the model name are required — gateways have no meaningful defaults, so a missing one is a loud, actionable error (serve logs it and disables /answer) instead of a silent dial to the wrong place. A base URL already ending in /v1 (the OpenAI SDK convention) works too. Leave the key unset (or set it to none) for keyless vLLM / llama.cpp deployments; embeddings are unaffected and keep their own provider. Redirects are refused outright — a gateway 3xx becomes a normal error instead of carrying the bearer token to another origin. Reasoning models pointed straight at the cloud OpenAI endpoint (o1 family) reject the classic fields; via the SDK, get_llm("openai", token_param="max_completion_tokens", options={"temperature": None}) renames the budget field and drops the fields they refuse (gateways normally translate this themselves).
Reasoning models (qwen3, deepseek-r1 and similar): mnemostack disables thinking by default (think=False in OllamaLLM) — with thinking on, these models spend the whole token budget on thoughts and return empty text, silently degrading reranking, expansion and extraction. Pass get_llm("ollama", think=None) to keep the model's own default, or think=True to force it on models that support thinking. Extra generation options go through options={...} (e.g. {"num_ctx": 8192}).
Run a local Qdrant for the vector store:
Optionally a Memgraph for the knowledge graph:
search and answer accept an optional --tier {1,2,3} flag that bounds how
much output a call produces. Useful when a recall is called from a long-running
agent loop where full recall output would burn context unnecessarily.
Omit --tier to get the full, uncapped output (backward compatible). Rule of
thumb for agents: tier 1 for navigation / existence checks, tier 3 only when
you actually need to read the memories. answer is already compressed, so it
needs a tier less often — use --tier 1 there to drop the SOURCES: block
when only the answer text is wanted.
When you want to feed items into mnemostack from code — a chatbot that logs every message, a scraper, a daemon tailing a log — use the Ingestor. It handles batching, deduplication, and idempotency for you.
Guarantees:
(source, offset, text). Re-running with the same input is a no-op: Qdrant upsert replaces the point onto itself, and an in-process LRU cache skips even the embedding call for items already seen in this session.batch_size, so provider HTTP overhead amortises across many items.indexed_at (UTC). Pass timestamp= (or metadata={"timestamp": ...}) to set the event time the temporal retriever filters on. With window_size > 1, sliding-window chunks also carry the window's temporal range as window_start_ts / window_end_ts payload keys.A memory stack that indexes only text answers "Not in memory" to questions whose answer lived in a photo. If your data contains images, describe them at ingest time and index the description:
describe_image is fully opt-in — nothing in the ingest or recall paths calls it, and text-only pipelines are unaffected. It works with any provider that has vision support (Gemini; Ollama vision models such as llava, llama3.2-vision, qwen2.5-vl) — providers without it return a normal fail-open error response. The default prompt produces a dense, index-oriented description (objects, any text/signs verbatim, setting, actions); pass prompt= to customize.
ing.stream(item_iter) yields per-batch stats so long feeds can be monitored without waiting for the whole stream to drain.failed but the rest of the batch still lands.Count and "list all X" questions need set completeness, which similarity top-K does not guarantee — the model counts what it sees and undercounts, returning a subset. For those, retrieve a wide candidate pool and enable the two-pass extract-and-aggregate mode:
list_extract_mode routes count/list questions through an extract pass (pulls every matching item as JSON) and a finalize pass (formats the list or count); other question categories are unaffected. The extract pass walks the whole pool you pass in, in batches of list_extract_batch_size (default 40), merging items across batches — so pool order does not decide whether a memory is seen, and the cost is one LLM call per batch plus finalize. An empty extract over a non-empty pool is retried once before abstaining. For guaranteed exhaustiveness on a bounded slice, build the pool from a full scan (VectorStore.scroll) filtered to the relevant slice. To evaluate it on your own data, the benchmark harness exposes the same knobs: benchmarks/locomo_single.py --list-extract --pool 150.
list_finalize="verbatim" skips the finalize LLM pass and assembles the answer deterministically from the extracted items (the count for count questions, the comma-joined items otherwise). Recommended for non-English corpora: an LLM finalize pass can paraphrase or distort items instead of repeating them verbatim. The default "llm" keeps the formatting pass.
Enriching payloads at ingest. Ingestor(enrich=callable) calls your function for every final item (including assembled window chunks) and merges the returned dict into the chunk payload — the mechanism is core, the extractor is yours (content extraction is corpus- and language-specific, so mnemostack ships none). Fail-open: a raising hook logs a warning and the item is indexed without enrichment; text/source/offset and an explicit item timestamp can't be overridden. From the CLI: mnemostack index docs/ --enrich mypkg.extractors:invoice_fields. Enriched fields combine with the rest of the stack: scope recall with filters={"amount": {"gte": 100}} and show them to the answer LLM with context_fields=["amount"].
Already-indexed collections don't need re-embedding to pick up enrichment: mnemostack index docs/ --enrich ... --refresh-payloads rewrites the payloads of existing chunks in place (Qdrant set_payload, vectors untouched) — only genuinely new chunks pay for embedding.
Structured payload fields in the answer prompt. By default the answer context shows each memory's timestamp, source and text. AnswerGenerator(context_fields=["author", "amount"]) additionally projects the named payload fields into each memory's context line (author=…, lists comma-joined, long values truncated; memories without the field render without it). Use it for structured facts the answer needs — who said it, amounts, your own ingest-time enrichments. Note the boundary: projection only changes what the answer prompt shows — retrieval ranks by text, so content that must be findable (image captions and similar) belongs in the text itself, not in a payload field.
Conversational follow-ups. "And who wrote that?" carries none of the conversation, so recall misses. rewrite_followup(query, history, llm) resolves pronouns and ellipses into a standalone question before recall — mnemostack holds no dialog state, you pass the history ((question, answer) pairs or plain lines, oldest first). One LLM call; the prompt instructs the model to return a self-contained question unchanged, and any failure falls back to the original query. To skip the call entirely for queries you already know are standalone, pass needs_rewrite=callable — that trigger heuristic is language-dependent, so core ships none (same boundary as question_classifier).
Non-English corpora. The built-in answer prompts and the question classifier are English; on other languages the extract/finalize passes degrade instead of helping. Both are pluggable:
Override names: the seven category prompts (general, list, count, temporal, multihop, inference, adversarial) plus list_extract / list_finalize. Required placeholders are validated at construction. mnemostack ships no translations by design — prompt quality is corpus- and domain-specific, so you own the templates.
This is the full runtime configuration. The LoCoMo numbers above are produced by a subset of it: the benchmark loop runs Vector + BM25 retrieval, the 8-stage pipeline, window_size=3, query expansion, and top-K 25 — the LLM reranker and the graph retriever are runtime-only features and are not part of the benchmark methodology (see benchmarks/run_locomo.sh for the exact reproduction path).
Reranker is generative: it asks an LLM to return candidate IDs. If you have
a backend that returns numeric relevance scores instead (a local cross-encoder
or a hosted rerank service), use ScoringReranker:
The scorer object only needs score(query, documents) -> Iterable[float].
Scores are relative; no absolute threshold is applied by default. Generative
LLMs can be wrapped as scorers, but dedicated rerank models/services are the
more stable default because they avoid ID-format parsing.
BM25Retriever needs a list of BM25Doc. Each doc is the atomic unit BM25 will rank — typically a paragraph or chunk of one of your source files:
For transcript-like inputs (adjacent user and assistant turns), prefer MessagePairChunker so related turns stay in the same chunk. See mnemostack.chunking.
If your canonical memory corpus is already stored in Qdrant payloads, build the BM25 corpus from the same collection instead of maintaining a separate markdown export. This keeps exact-token lookup aligned with vector search (IDs, commit hashes, filenames, quoted phrases):
You can also call bm25_docs_from_qdrant(...) directly if you want to combine Qdrant payload chunks with local BM25Docs before constructing BM25Retriever.
For morphologically rich languages or domain-specific normalization, pass a custom tokenizer/analyzer. The same analyzer is applied to corpus and query text; the default exact-token behavior is unchanged when omitted.
If you pre-tokenize BM25Doc objects yourself, pass retokenize=False when
constructing BM25/BM25Retriever with the same analyzer. The
BM25Retriever.from_qdrant(...) helper does this automatically.
If you want mnemostack available to callers that aren't Python — any service written in Node, Go, Rust, or a plain curl from a shell script — install the server extra and expose it over HTTP:
mnemostack serve binds to 127.0.0.1 by default. Use
--host 0.0.0.0 only behind your own auth/rate-limit layer.
Endpoints:
| Method | Path | Purpose |
|---|---|---|
GET | /health | Qdrant + Memgraph reachability + config summary |
GET | /healthz | Liveness probe — 200 whenever the process is up (no backend checks) |
GET | /readyz | Readiness probe — 503 when Qdrant is unreachable (graph is fail-soft, never gates) |
GET | /status | Operator snapshot — config, live dependency reachability, headline counters |
POST | /recall | Hybrid recall with optional 8-stage pipeline |
POST | /answer | Recall + LLM answer synthesis with citations |
GET | /resolve/{chunk_id} | Verify a citation — resolve a chunk id back to its source document — read |
POST | /feedback | Explicit click/usefulness feedback for stateful learning |
POST | /memories | Create memories (server-side embedding, store-backed dedup) — write |
GET | /memories | List what the tenant holds from one source (ids + integrity metadata, no text) — read |
DELETE | /memories | Irreversible erasure by id list or source — write |
POST | /invalidate | Non-destructive retraction by id list or source — write |
POST | /triples | Write knowledge-graph facts — write |
GET | /metrics | Prometheus scrape endpoint (counters + summary histograms) |
GET | /docs | Interactive OpenAPI UI |
Response shape (abridged):
The order of results is authoritative — do not re-sort by score: many stages and fallback paths write that number on different scales, and a rerank changes the order without rewriting it. See what score is not.
Pass "include_trace": true in the request body to additionally get a trace object with per-retriever ranked lists, the fused order, and the post-rerank order — useful when debugging why a memory did or didn't surface.
Pass "token_budget": 2000 to cap how much prompt space the results may occupy: the final ranking is cut to the prefix whose total text tokens fit the budget (a hard cap — never overshot, so an oversized top hit yields an empty list rather than a blown prompt). tokens_estimate in the response is the value the budget is enforced against; counting uses a dependency-free heuristic (≈4 chars/token for ASCII, ≈2 for non-ASCII scripts), so leave yourself margin rather than budgeting to the exact context limit. A server-wide default can be set with recall.token_budget in the config file (or MNEMOSTACK_TOKEN_BUDGET); per-request values override it. The same parameter is available on /answer (caps the memories fed to the LLM), on MCP mnemostack_search / mnemostack_answer, on the CLI as --token-budget, and in the library as recall_flow(..., token_budget=...) — where you can also pass an exact token_counter= (e.g. a tiktoken encoder) instead of the heuristic.
Pass "filters": {...} to scope recall by payload fields — exact match ({"tenant": "a"}) or inclusive ranges ({"timestamp": {"gte": "2026-01-01"}}). Filters apply inside every retriever, not as a post-filter on the output: the candidate pool itself is restricted, so top-K stays full and results never include points outside the scope — this is the isolation contract for multi-tenant and per-user memory. Sources that cannot attribute their results to the scope contribute nothing rather than leak. The knowledge-graph retriever attributes its hits where it can: a filter key the hit's own node metadata carries (e.g. index_root) is checked in place, the rest is proven through the hit's vector chunks — a graph file hit (with a recorded root, pinning the probe to its exact document) passes the filter exactly when at least one of its chunks does; entity nodes and anything else without pinnable chunks are still excluded, never leaked. The same filters parameter is available on /answer (the answer is generated only from in-scope memories, including retry sub-recalls), on MCP mnemostack_search / mnemostack_answer, on the CLI as --filters '{"tenant": "a"}', and in the library as recaller.recall(query, filters=...) / recall_flow(..., filters=...).
The /answer endpoint adds { answer, confidence, sources } alongside the memories and carries the same degraded / notes / opt-in trace fields, plus tokens_used — the LLM provider's reported token usage for the generation call that produced the answer (provider-specific semantics; null when the provider reports nothing). If the LLM isn't configured, /answer returns 503 and /recall still works — graceful degradation applies at the HTTP layer too.
Start the server with --retry-on-weak if you want a recall that comes back nearly empty to be paraphrased by the answer LLM and asked again, fusing the rounds by reciprocal rank — so a memory that two phrasings both find outranks one that only a single phrasing did, and a later paraphrase can beat an earlier one. What counts as "weak" is a COUNT (--retry-weak-below, default 1 — only a recall that returned nothing), not a score: fused scores are RRF values encoding rank, not confidence, so a threshold on them would measure nothing. One extra round, at most two paraphrases, and every retry carries the caller's tenant, filters, validity view and budget unchanged. Budget for it accordingly: each variant repeats your recall in full, reranker included, so a weak recall costs up to three LLM calls (one paraphrase plus one rerank per variant) and two extra retrieval rounds — not the single call the name suggests. Two consequences worth knowing before you switch it on: a retried response fuses over several rounds where an unretried one fuses over a single one, so the rank basis behind score differs and the two are never comparable — threshold on rank, not on the number, and see what score is not for what the number holds on each path; and a retry never returns fewer memories than it was given — if the fused list plus your token budget would hand back less than the recall already had, the original results come back untouched. That is a guarantee about the size of the response, not its membership: at a fixed limit a hit that a paraphrase ranks first will take the slot of one your original phrasing ranked last, which is what asking again is for. Off by default, and a request can only opt out of it — the server pays for the LLM call, so enabling it is the operator's call.
Stateful learning is explicit. Start the server with --auto-record-ior if you want /recall and /answer responses to update inhibition-of-return state, and with --record-access (or MNEMOSTACK_RECORD_ACCESS) if you want them to stamp access_count/last_accessed on every point they return — the reinforcement the freshness stage reads, recorded where the retrieval actually happens instead of in each client. Both are off by default: they turn reads into writes. Access recording is fail-open (a failed write is logged and counted, never raised), best-effort on the count (no atomic increment exists; concurrent recalls of one point can record one increment, and the reader clamps reinforcement at 10 anyway), and scoped to the caller's tenant.
It also changes what ranking means, so know which direction it moves things: recording these keys makes the freshness stage's access term live, and that term is a bonus, not a decay. A memory that has been retrieved gets a multiplier in [1.0, 1 + access_bonus_max] (0.25 by default, reached at 10 accesses) which fades back toward 1.0 as the access ages and stops there — use can raise a memory's rank, never lower it. A memory nothing has ever retrieved sits at exactly 1.0, so a deployment that records no accesses ranks exactly as it did before. Set --access-bonus-max 0 (or MNEMOSTACK_ACCESS_BONUS_MAX=0, honoured by serve, search, answer and mcp-serve alike) to take the access signal out of ranking entirely — the switch to reach for if your clients stamp last_accessed themselves, since leaving --record-access off does not help there: the stage reads those keys whoever wrote them. Values are clamped to [0, 1]. Because this is the one term a recall's own output feeds back into, it is bounded on purpose — small ceiling, saturating counter, and only the points actually handed to the caller are recorded. Send user actions to /feedback to update Q-learning:
signal is one of useful, clicked, or irrelevant; pass the retrievers list returned by /recall as sources so Q-learning can update the right source weights.
The same state update is available from CLI as mnemostack feedback ... and from MCP as mnemostack_feedback.
Multi-tenant auth. By default the server is unauthenticated (single-tenant; put it behind your own auth layer). For a hard, per-tenant boundary, start it with --auth and issue service keys:
The tenant is resolved from the key (a client can't assert another's), enforced across the vector store, the tenant-scoped knowledge graph, and per-tenant learning state — so it's a real authorization boundary, unlike the caller-supplied filters above. /recall, /answer and GET /memories require read; /feedback, POST/DELETE /memories, /invalidate and /triples require write; a missing/invalid key is 401, insufficient scope 403. Cap each tenant with mnemostack quota set --tenant <id> --max-points N --max-rps R (storage enforced at ingest, rate on the HTTP surface → 429). The operator endpoints (/health, /healthz, /readyz, /status, /metrics) stay unauthenticated — protect them at your proxy if sensitive. Auth is off by default; you can still front the server with your own reverse proxy (nginx, Caddy, Traefik) either way. See docs/deployment.md and docs/api-stability.md.
Current graph facts use the explicit valid_until="current" marker. If you
created graph data with an older release, run
mnemostack graph-migrate-current --dry-run first, then
mnemostack graph-migrate-current to backfill legacy NULL markers.
Add the entry below to your Claude Desktop config file:
~/Library/Application Support/Claude/claude_desktop_config.json%APPDATA%\Claude\claude_desktop_config.json~/.config/Claude/claude_desktop_config.jsonClaude will then be able to call mnemostack_search, mnemostack_answer, and graph tools.
embed/embed_batch are the neutral primitives and stay the required
surface. Ingestion and retrieval call the inherited role methods —
embed_document(s) for chunks, embed_query/embed_queries for queries —
which apply the model family's embedding profile exactly once before
delegating to your primitives. Built-in profiles cover asymmetric families
(Qwen3-Embedding's query instruction; E5's query: /passage: prefixes;
the multilingual-E5-instruct variant's instruction format);
symmetric models resolve to the identity profile and behave exactly as
before. Transforms touch only inference input — stored text, chunk ids,
lexical search and citations keep the original document text. To teach
mnemostack a new family without patching it:
Indexed points are stamped with a document-space fingerprint (provider +
model + document transform + dimension + provider inference knobs —
profile name/version are metadata and never hashed; upgrading from 1.x? see
docs/migration-1.x-to-2.0.md); mnemostack index /
index-markdown refuse a collection embedded under a different space, and
doctor reports the active profile and fingerprint.
If you run an OpenClaw-style assistant with its own scripts/recall.sh or similar recall entry point, you don't have to wire mnemostack in by hand. Ask the assistant to do it and it can follow the steps below verbatim.
(These steps are written against the markdown-backed memory layout popularised by OpenClaw / OpenClaw — the reasoning applies to any assistant whose memory lives in MEMORY.md + daily notes.)
Assistant, read this section and follow the steps if your user asks you to install mnemostack.
Install the package into whichever Python env your recall scripts use.
Confirm with python3 -c "import mnemostack; print(mnemostack.__version__)".
Make sure Qdrant and (optionally) Memgraph are reachable. If either is missing, start them:
Check with curl -sf http://localhost:6333/collections and echo 'RETURN 1;' | cypher-shell -a bolt://localhost:7687.
Populate the vector store with the assistant's existing memory. Almost every OpenClaw-style assistant stores memory as markdown files — MEMORY.md, memory/*.md daily notes, maybe USER.md / TOOLS.md. Those files are the ground truth; Qdrant is just a searchable index on top of them. So "migration" here really means a first-time ingest.
Two cases:
Case A — user has no Qdrant yet (default OpenClaw setup, markdown-only memory).
Stand up Qdrant (see step 2), then index the workspace in one command:
This walks every *.md / *.txt under the workspace, chunks them, embeds each chunk with the chosen provider, and upserts into Qdrant. For a typical workspace (a few megabytes of notes) this finishes in a minute or two and stays well inside free-tier embedding quotas.
Re-running is safe. mnemostack index uses a deterministic id per (source, offset, content) chunk, so a second run skips unchanged chunks (no duplicates, no wasted embedding calls) and only re-embeds chunks whose content changed or moved. Use --recreate if you deliberately want to wipe and rebuild from scratch.
Memgraph is optional — skip it for a pure markdown-backed memory and MemgraphRetriever just won't be part of the retriever list.
Case B — user already has Qdrant / Memgraph populated by another plugin or a prior assistant run. Don't re-index. mnemostack reads existing collections and graphs as-is; payload like chat_id, memory_class, valid_from is preserved. Just point VectorStore(collection=...) at the existing collection name and MemgraphRetriever(uri=...) at the existing graph.
Critical for both cases: use the same embedding model for the query that was used for the stored points. Mixing providers/models across ingest and query silently returns garbage — no error, just bad results. Indexing commands now enforce this for documents: points carry a document-space fingerprint and index/index-markdown refuse a collection embedded under a different provider/model/profile.
Point your recall entry at mnemostack. Swap your search helper to use the 4-retriever Recaller from the "Full stack" example above. If your recall script has a --unified flag, route to mnemostack first and fall back to the legacy code path on any exception.
Keep the fallback wired. Don't delete the existing recall implementation. Route mnemostack as the primary, fall back to the old recall on any exception. A bad upgrade becomes a one-minute rollback, not an outage.
Smoke-test on real queries from the user's memory, not toy strings. Ask at least one exact-token query (an IP, a ticker, an error string), one semantic query, one temporal query ("what did we decide last March"), and one list query. If any regresses vs the old recall, keep both running in parallel until you understand why.
Write a short note in the user's daily memory describing what changed: which command was used to ingest, which embedding provider, which collection name, and where the fallback lives. Future sessions need this to avoid re-indexing with a different model and trashing recall quality.
If you're a human operator and you'd rather do this yourself, the same steps apply; treat them as a checklist.
This repository also includes a self-contained OpenClaw companion plugin in openclaw-plugin/. mnemostack remains the recall backend — vector + BM25 + graph + temporal retrieval, fusion, reranking, and answer synthesis — while the plugin connects that backend to OpenClaw's before_prompt_build hook.
Zero-config path: install mnemostack, run the daemon on the default local port, install/enable the plugin, and OpenClaw will automatically inject bounded recall answers for recall-style questions:
The plugin defaults to http://127.0.0.1:18793/answer, supports English/Russian trigger defaults with extensible language-agnostic trigger lists, and can fall back to a Script backend such as recall-selfeval.sh when you are not running the daemon.
mnemostack health/doctor/inspect/search/answer/index/mcp-serve)mnemostack.graph.TripleExtractor)mnemostack.config, mnemostack init/config CLI)mnemostack.vector.AsyncQdrantStore)examples/docker-compose.yml)benchmarks/run_locomo.sh)pip install 'mnemostack[server]', mnemostack serve)Recaller.recall_async and parallel retriever dispatch (proven: 5 concurrent HTTP recalls complete in ~1x single-request wall-clock)benchmarks/synthetic_longhorizon.py)Ingestor API (mnemostack.ingest)/metrics endpoint on the HTTP serverMemgraphRetriever probes (telegram_id, handle, name_lower)/metrics (mnemostack_recall_<name>_latency_ms)reciprocal_rank_fusion(weights=[...]))Recaller (adaptive_weights=True)HyDERetriever (opt-in, not in default Recaller)search/answer (--tier {1,2,3}, backward-compatible)integrations/)TemporalRetriever + dispatch-by-type filter builder (0.2.0a1)recall_flow() — CLI/HTTP/MCP rank identically (0.5.0)index --prune (root-scoped) + --refresh-payloads without re-embeddingfilters= on every surface with adversarially-tested tenant isolationIngestor(enrich=...)) + context_fields answer projectionthink control (off by default) + generation options passthroughrewrite_followup)Issues and PRs are welcome. Public APIs are intended to remain stable; new functionality should land additively where possible.
Apache 2.0 — see LICENSE.