Persistent agent memory with a decision layer: replay, restore, verify, or none β explainable.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
One-click editor setup isnβt available for this listing yet β we donβt have a confirmed install command, and weβd rather show nothing than point your editor at the wrong package or host. Follow the projectβs own setup instructions, linked above.
Persistent semantic memory for AI agents with intelligent decision-making.

π Created by: TheProdSDE
Most AI memory systems retrieve and inject past context into every prompt. This leads to wasted tokens, inconsistent responses, and agents that blindly replay stale or wrong answers.
Agent Memory adds a decision layer:
Every resolve() returns an explicit action with a scored, explainable rationale β
not just a retrieved chunk.
Adversarial eval: 34/36 (94%) on trap queries β the 2 misses return VERIFY (cautious), never a wrong REPLAY β see benchmarks.
Context rot is the measured degradation of LLM accuracy as the context window fills β long before the token limit. Stale chunks, irrelevant retrievals, and unbounded conversation history don't just waste tokens; they actively degrade answers ("lost in the middle", instruction drift, distractor sensitivity).
Context rot has two causes. Agent Memory addresses the first; nothing outside the model itself can address the second.
1. What goes into the context β controllable, and this SDK's job:
| Rot source | Mechanism in Agent Memory |
|---|---|
| Irrelevant memory injected into every prompt | Decision layer β NONE refuses to inject when nothing truly matches (34/36 on adversarial trap queries; the 2 misses fail safe to VERIFY) |
| Unbounded in-session history | PagedMemory β fixed in-context buffer; old turns page out to recall storage and return per-query (MemGPT-style tiers) |
| Instruction drift in long coding sessions | RESTORE re-injects the relevant convention fresh, near the end of context, exactly when a query needs it |
| Stale facts silently reused | VERIFY + custom verifier callbacks + TTL expiry + half-life temporal decay |
| Reminders that never adapt | mark_correct() / mark_wrong() β confidence learning promotes memories that keep helping, demotes corrected ones |
| Knowledge lost when the session ends | from_conversation() distills durable facts from conversation turns into the store |
2. How the model attends over tokens already in its context β not controllable from outside. Attention degradation over long context is a property of the model. No memory layer changes that. What Agent Memory does is keep the context small and relevant enough that the model rarely enters the degraded regime in the first place.
The honest claim: Agent Memory prevents context pollution β the dominant
controllable cause of context rot in agentic systems. It doesn't change model attention
behavior, and it only helps if your agent routes context through resolve() /
PagedMemory instead of concatenating history by hand.
Today you need three tools wired together to get what resolve() does in one call:
a semantic cache (GPTCache) for replay, a memory layer (Mem0 / Zep) for context, and
custom staleness logic for verification. No existing tool decides β at read time β
whether and how a memory should be used.
Cells about other projects are capability checks against their own source or
docs (mem0 checked on 2026-09-27), not measured behaviour; β means we have not
tested it rather than that it is absent. Corrections welcome β see
docs/comparison.md for sourcing and the
benchmark RFC for the review process.
| Capability | Mem0 | Zep / Graphiti | Letta (MemGPT) | GPTCache | Agent Memory |
|---|---|---|---|---|---|
| Read-time decision (replay / inject / verify / skip) | β always injects | β always injects | β οΈ LLM self-manages | β οΈ replay only | β REPLAY / RESTORE / VERIFY / NONE |
| Explainable per-decision scores | β | β | β | β | β
decision.explain() |
| Semantic answer cache (skip the LLM call) | β | β | β | β | β |
| Staleness protection at read time | β οΈ write-side updates + expiration_date | β temporal graph | β | β οΈ eviction only | β VERIFY + TTL + confidence decay |
| Adversarial trap-query eval published | β none found | β none found | β none found | β none found | β 34/36 (94%) |
| LLM / API calls per memory op | 1+ | 1+ | 1+ | 0 | 0 |
| Local after model assets are installed/cached, zero API keys | β οΈ self-hostable; needs an LLM for extraction | β οΈ needs server + LLM | β οΈ LLM per op | β | β SQLite + local ONNX |
| Paged context tiers (MemGPT-style) | β | β | β | β | β
memory.paged() |
Because hosted or model-backed configurations can add an LLM or embedding API round-trip per memory operation, their latency includes provider, model, and network costs. Exact latency depends on each project's configuration; this repository does not publish a universal 100msβ2s floor. Agent Memory resolves in-process in SQLite-only mode; measured latency depends on corpus shape and cache state (see performance and stress-testing details).
On retrieval, we publish a cleaned-release retrieval-proxy measurement below. On end-to-end accuracy (LLM answering + judge, where Mem0 and Zep publish), we don't quote numbers we haven't measured yet β that stage is next on the roadmap. β Full feature matrix and trade-offs (including where they're better): docs/comparison.md
LongMemEval is a benchmark for conversational-history retrieval. These results measure Agent Memory's RESTORE/retrieval tier, not REPLAY, VERIFY, TTL, or end-to-end answer correctness. Each question runs against a separate SQLite store containing that question's haystack; aggregate ingestion totals are not the size of one queried store. All ingestion uses zero LLM calls and $0 in API charges:
LongMemEval_S (500 independent ~48-session haystacks; 124K turn-pair entries across all runs): 98.1% session Recall@5 with local ONNX embeddings, 96.0% lexical-only, 10.04ms lexical / 19.56ms semantic p50 retrieval. The report includes p90/p95/p99 and run-resource measurements.
LongMemEval_M (500 independent ~500-session haystacks; ~2,500 turn-pair entries per queried store): 87.0% session Recall@5 with lexical retrieval, 12.05ms p50. This uses the official cleaned re-release and turn-pair indexing; it is not directly comparable with the paper's original-release session-index baselines.


Full methodology, per-type tables, scope notes (what this benchmark does and
doesn't test), and negative results are in the
benchmark report. Reproduce the semantic
_S result with uv run python benchmarks/longmemeval/run_retrieval.py --semantic.
Every claim below is reproducible from this repo:
agent-memory eval (methodology)No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/agent-memory-5)<a href="https://allmcps.com/mcp/agent-memory-5"><img src="https://allmcps.com/api/badge/agent-memory-5?style=directory" alt="Agent Memory on AllMCPs" /></a>