The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Memshelf MCP listing page.
Put your agent's memory on a shelf, hand it the index.
Long-running agent sessions burn tokens re-sending history and lose detail to lossy auto-compaction. memshelf applies the docshelf pattern — tiny index in context, bodies fetched on demand — to the agent's own working memory:
INDEX.md (kilobytes) + digests in context and
recalls exact sections via INDEX → episode → section slice over MCP.Positioning in one sentence: claude-mem's loop, git's substrate, docshelf's navigation — episodic memory you can grep, diff, review, and carry between hosts. Private and local by default: the standard storage mode is a local git repo with no remote configured. The tool is public; the memory never is.
One week of dogfooding on the live shelf — full numbers and methodology in
docs/demo.md:
| Measure | Result |
|---|---|
| Episodes on the shelf | 34 |
| Standing cost in every session (INDEX + digests) | ~8.6K tokens |
| Shelved mass those episodes replace | ~1.9M tokens — ≈220 : 1 |
| One question answered from memory | ~1.8K tokens (INDEX + one episode) |
| Recall test: fresh agent, INDEX path only | 5 / 5 — zero misses, zero over-fetch |
Tokens are counted as chars/4 everywhere, so the ratios are estimator-independent; absolute counts move with the tokenizer.
As an MCP server:
As a Claude Desktop extension — adapters/claude-desktop/:
an .mcpb bundle installed from Settings → Extensions, with a Default
shelf setting so calls need not repeat the path. Nothing has to be installed
alongside it — not even Python.
As a Claude Code plugin — adapters/claude-code/:
a /shelve skill plus SessionStart / SessionEnd / PreCompact hooks.
Or from the shell (pip install memshelf-mcp, Python ≥ 3.10) — the same
loop, no MCP:
pip install 'memshelf-mcp[semantic]' adds an embedding sidecar so search
also finds paraphrases and the other language (memshelf semantic build --shelf ~/my-shelf once; MEMSHELF_SEMANTIC=off turns it off). Optional:
the base install stays grep-only and as light as it is.
One verb per job; the same names over MCP (memshelf_*) and in the CLI
(long-form descriptions: docs/tools.md):
| Tool | What it does |
|---|---|
init | Create (or top up) a memory shelf: docshelf layout, fixed categories |
shelve | Offload one closed topic as a durable, indexed episode; --amend rewrites in place |
lint_digest | Validate a digest against the contract without touching the shelf |
import | Retro-shelve a whole exported dialog without pulling it through context |
index | Return the shelf INDEX — the small recall entry point |
recall | Fetch an episode by id, or a single ## Section of it |
search | Grep the shelf; returns matching episodes |
stats | The shelf's token economy: standing cost vs shelved mass, claimed vs realized |
advise | What your context is made of and what you could put down — proposals only |
rebuild | Regenerate every derived file from the episodes |
rollup | Archive a period behind one digest-of-digests |
purge | Drop episodes past retain_until, then reindex — dry run by default |
resolve | Settle multi-writer conflicts: regenerate derived, union the recall log |
doctor | Diagnose: episode schema, digest contract at rest, secret shapes, index bloat |
prune-splits | CLI only — remove H2 split directories git never got (migration for #109) |
tags | CLI only — episodes grouped by frontmatter tag (#18) |
graph | CLI only — who mentions whom: cross-episode id references as JSON or Mermaid (#18) |
retro | CLI only — one quarter of the shelf as a Markdown retrospective (#18) |
fork | CLI only — bootstrap a fresh session from INDEX + selected episodes or sections (#18) |
mirror | CLI only — INDEX (± episodes) as one self-contained HTML page for phone-side reading (#18) |
semantic | CLI only — build / status / drop the embedding sidecar that turns search hybrid; lives outside the shelf, needs pip install 'memshelf-mcp[semantic]' (#17) |
search-bench | CLI only — hit@1 / hit@k / MRR of grep vs hybrid on a query<TAB>expected-id file (#17) |
The digest is a contract, not a convention. It is the only thing read at
recall before fetching a body, so a weak one devalues the whole episode.
lint_digest runs the same validator as shelve with no side effects
(--strict turns warnings into failures); errors block a shelve, warnings do
not — a pure reference digest legitimately carries no decision marker. A
rejected digest is a feature: the tool prints exactly what to fix and writes
nothing.
--amend re-runs the whole pipeline — redaction, the digest contract,
composition — so an amended episode is exactly as guarded as a fresh one,
which a hand-edit of the file never is. Amending a slug that is not on the
shelf is an error, not a create.
The episode is the source; everything else is output. ledger.tsv,
INDEX.md, stats.svg and each category's .meta.json are derived:
shelve writes and commits the episode alone, rebuild renders the rest —
delete all four and rebuild restores them byte-identically. That is what
makes two sessions shelving in parallel a non-event: the merge is clean by
construction. On a shared shelf, let a bot own the derived files on main —
ready-to-copy workflows in adapters/shelf-repo/;
rebuild --adopt migrates an older shelf once, rebuild --check is the
CI guard.
Two consequences worth stating plainly, because getting them wrong costs a merge conflict:
doctor reports no-ledger-row and stale-index immediately after a
correct shelve — on every branch, main included. Nothing is broken:
the episode is written, the derived files are not rendered yet. They clear on
the next rebuild — the bot's run, on a shelf that has one.ledger.tsv/INDEX.md/stats.svg meets the bot's, and the merge stops
being clean by construction. Wait for the renderer; on a shelf without a bot,
run memshelf rebuild --shelf . as its own step.If those warnings persist for a day while episodes keep arriving, that is a
different state — the renderer is not lagging, it is stopped — and doctor
says so separately, as derived-stale at error severity. The day is counted
from when the renderer could first see the work, not from the ledger's last
commit: an episode pushed minutes ago onto a shelf whose ledger has not moved
since yesterday says nothing about the renderer, and saying otherwise sent
readers to a manual rebuild, which is the conflict this whole split exists
to avoid.
That arrival is read from this clone's reflog for the tracked upstream — the
one local record of when the ref moved here. A commit date is not a
substitute: it says when the episode was written, and «shelve now, push when
confirmed» is a documented way to work, so the two can be a working day apart.
Where the reflog cannot say — a fresh clone starts an empty one, which is what
CI and ephemeral agent sessions run in — doctor reports
renderer-wait-unknown at the unknown level instead of picking a verdict:
from there, «stopped» and «handed the work a minute ago» look the same, and
the renderer has to be judged where it can be observed, on its own job's run.
On a shelf with no upstream there is no renderer to be fair to, and the
ledger's own age stays the clock.
There was a third way to hold stale-index forever, and it is fixed rather
than documented: docshelf split any episode past 50 KiB into section files
beside it, shelve committed the episode alone, and from then on this working
copy rendered an INDEX no other checkout could produce — no rebuild could
clear it, and search answered with addresses that existed on one machine
(#109). shelve no longer splits. A shelf that already carries such
directories keeps reporting them as local-split-dir until
memshelf prune-splits --shelf . --apply removes them; the episode file holds
every section, so nothing is lost. It is a dry run without --apply, and a
split directory that is committed is reported and left alone.
advise proposes, never writes. It answers the question the project was
founded on — a dead topic has been occupying 30K tokens for forty minutes.
The tool cannot see your window, so you tell it what is in there:
Three things keep it honest: it counts itself (INDEX + digests are in the
report, not left out of it), it verifies episode= claims before
proposing a drop, and it reports net — a topic too small to pay for its
own digest is not proposed at all.
Rollup shrinks navigation and nothing else. When INDEX grows into a real
share of your window, rollup collapses a period into one digest-of-digests
and moves the originals to archive/ — still reachable by recall and
search, every ledger row intact. The rollup digest is yours, not the tool's:
synthesizing a quarter is the part a tool cannot do.
index-bloat is not what a rollup is for. INDEX lists your episodes, so
its size grows with the shelf by design; its budget grows with the shelf too
(INDEX_BASE_TOKENS + INDEX_TOKENS_PER_ENTRY × listed). Over budget therefore
means entries are overpriced, never that there are too many of them — so
doctor reports the cost of one line, and the fix is to trim it and
rebuild. A rollup would remove entries and their allowance together and
leave the price where it was. Having the two paired the other way is what made
"archive a third of your memory" the standard way to silence a formatting
problem.
purge deletes the working tree, not history. Retention is opt-in per
episode (--retain-until); purge is a dry run until --apply — and even
then git history still has the file. Real erasure is a deliberate
filter-repo pass over the whole repository, never a side effect of a tool
call, and the purge report says so.
resolve regenerates derived paths, never merges them — a derived file
has no history, only a current correct value. The one file it unions is
recall-log.tsv, because a recall is an event, not a fact about the
episodes. Conflicting episodes are content, not mechanics: resolve
reports them and steps aside.
The design rationale behind each rule lives in
docs/DECISIONS.md and
docs/ARCHITECTURE.md.
The memory is vendor-portable, and that is a measured fact, not a design
intention: the same live shelf has been read and cross-written by Claude Code
(Anthropic) and Gemini CLI (Google) through one shelf-spec server —
protocol and field notes in docs/portability.md.
M0 complete: the pattern was validated with zero code on a live shelf —
retro-import of months of material, then a week of shelve-at-close
(docs/M0.md). M1 shipped the server/CLI that enforces it,
plus the Claude Code plugin. Next milestones with exit criteria:
docs/ROADMAP.md; release history:
CHANGELOG.md.
Rendered site: https://ignatenkofi.github.io/memshelf-mcp/ — including the week-report infographic from the dogfood shelf.
| Doc | What it covers |
|---|---|
docs/MANIFEST.md | Problem, the bet, hero scenarios, principles, non-goals |
docs/tools.md | Tool reference: the long-form description of every memshelf_* tool (the MCP schema carries only the short one) |
docs/ARCHITECTURE.md | Episode format, digest contract, storage modes, triggers, MCP tool surface, portability model, privacy, failure modes |
docs/LANDSCAPE.md | Prior-art survey (2026-07), platform built-ins, positioning, risks |
docs/ROADMAP.md | Milestones M0–M3 with exit criteria |
docs/DECISIONS.md | Decision log |
docs/M0.md | M0 experiment protocol and results: cases, token ledger, recall test |
docs/demo.md | Measured numbers from the dogfood shelf: compression, recall test, doctor findings |
docs/portability.md | One memory, multiple AIs: the cross-vendor experiment |
docs/examples/ | A worked episode file and a memory-shelf INDEX |
adapters/claude-code/ | Claude Code plugin: /shelve skill + SessionStart/SessionEnd/PreCompact hooks |
adapters/claude-desktop/ | Claude Desktop .mcpb extension: builder, bundle checker, default-shelf setting |
shelf/ is this project's own memory shelf, dogfooding the tool it ships:
session episodes with decisions, rejected options and their reasons. Before
asserting anything about past work here, read shelf/INDEX.md
and fetch the one episode it points at — a guess about a past decision is a
defect, not an estimate. Derived files (ledger.tsv, INDEX.md, stats.svg,
.meta.json) are written by the same commit as the episode
(memshelf rebuild --shelf shelf); the shelf-pr-guard workflow fails a PR or
a push to main where they drifted.
Designed as RFC-0001 in the docshelf-mcp repo (#42, #43, #44); this repo is the project's home from 2026-07-13 on. The docshelf copy is frozen as a historical snapshot.
MIT — see LICENSE.
mcp-name: io.github.ignatenkofi/memshelf-mcp