The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the MemPersist listing page.
Mempersist is a clean-room, Cloudflare-native long-term memory service for AI conversations. It preserves original ChatGPT exports and normalized conversation graphs in R2, catalogs them in D1, builds disposable lexical and semantic indexes, and exposes compact retrieval and intentional writes through MCP.
It does not silently capture ChatGPT traffic, extract replacement “facts,” provide a SaaS billing layer, or make search indexes canonical. “Unlimited” means no application message quota; Cloudflare limits and billing still apply.
R2 is the source of truth. D1 holds operational metadata and the derived FTS representation. Vectorize is disposable. A canonical write succeeds before indexing is queued, and an indexing failure never reports that durable memory was lost.
See ARCHITECTURE.md, SECURITY.md, and docs/operations-and-recovery.md.
yarn wrangler whoamiUse Yarn only.
Set a long random MEMORY_API_TOKEN in .dev.vars. Local D1 and R2 are simulated; Workers AI and Vectorize bindings are remote in the main configuration. Unit and integration tests do not call remote AI.
Provisioning is intentionally manual and must be explicitly authorized. Follow docs/cloudflare-resources.md, then add the real D1 database_id returned by Wrangler to wrangler.jsonc. Never invent IDs or reuse unrelated account resources.
Set the production secret without putting it in source:
Apply migrations and deploy only after review:
Export data from ChatGPT, extract conversations.json, then:
Files up to 16 MiB use direct streaming upload. Larger files use 16 MiB R2 multipart parts — the
deployed max_direct_import_bytes and max_multipart_part_bytes values from
memory_get_capabilities. The Worker hashes the completed object, preserves it unchanged, detects
exact duplicate exports, and processes at most 25 conversations per queue turn. Check progress with:
The Streamable HTTP endpoint is https://<worker>/mcp. Interactive clients such as ChatGPT
use OAuth 2.1 authorization-code flow with PKCE: the consent page takes an email, sends a
single-use magic link, and completes the connection only after the link is opened. An existing
email reconnects to its archive; a new archive is created after the first link. Each account can
own multiple namespaces, and the same namespace name may exist in different accounts with fully
separated data.
Developer scripts and the CLI may keep sending MEMORY_API_TOKEN as a bearer token for the
owner archive.
Email authentication uses the EMAIL send binding. The sending address per endpoint is
configuration (AUTH_EMAIL_FROM, LEGACY_AUTH_EMAIL_FROM), not part of the API contract.
The public site, OAuth pages, and magic-link email support English and Bahasa Indonesia. Use the
language switcher to persist a browser preference; otherwise MemPersist uses Accept-Language and
falls back to English. API, MCP, and CLI contracts remain English.
Open /login for passwordless access to the server-rendered dashboard. It includes archive totals,
a deterministic memory map, revision-pinned canonical conversation reading, display-name editing,
and a streamed lossless export of current memory. Namespace emptying runs asynchronously. Account
deletion has a seven-day cancelable grace period and makes writes read-only while pending.
See docs/dashboard.md and ADR 0028. No frontend framework, extra dependency, or additional Cloudflare resource is required.
For the deployed Worker, add https://mempersist.codifiedtech.id/mcp as a custom MCP app in
ChatGPT Developer mode. ChatGPT discovers OAuth automatically, opens the consent page, and
stores the issued access/refresh tokens. Existing connections keep working after upgrades
without re-authorization. Clients already configured with the legacy endpoint remain
supported; changing one to the primary endpoint requires one new authorization. Do not paste
MEMORY_API_TOKEN into ChatGPT's connector settings.
Available tools:
memory_searchmemory_get_contextmemory_get_conversationmemory_get_conversationsmemory_list_conversationsmemory_list_revisionsmemory_resolve_conversationsmemory_build_contextmemory_list_namespacesmemory_statsmemory_storememory_appendmemory_replacememory_restore_revisionmemory_copy_conversationsmemory_get_capabilitiesmemory_store and memory_append accept optional tags (lowercased, deduplicated, up to 20);
memory_search filters by tags with AND semantics and returns each conversation's tags. See
docs/mcp.md and ADR 0013.
memory_delete_conversationsmemory_empty_namespacememory_import_statusSearch returns compact references; call memory_get_context only for selected results. See docs/mcp.md.
memory_get_capabilities returns the deployed capability contract: protocol and capabilities
versions, transport and message limits, per-tool item and byte bounds, and feature flags. The
limits quoted below are the same enforced values it reports, so treat that tool as the single
contract rather than independent figures. Aggregate limits are measured as UTF-8 bytes of the
complete serialized request before any canonical write begins: inline JSON writes on both
transports are limited to 1,048,576 bytes (1 MiB), and imports are limited to 16 MiB direct and
16 MiB per multipart part. An oversized request fails with code REQUEST_TOO_LARGE and a
conservative suggested_max_items estimate, never a silent truncation.
For known memories, start memory_get_conversations with a requests array of 1–20
conversation requests and optional max_serialized_bytes (default 32,768; minimum 4,096;
maximum 49,152 — the deployed values memory_get_capabilities reports). A continuation sends
exactly one opaque cursor plus the optional budget —
never both requests and cursor, and never neither. Results use the established camelCase
batch fields batchId, results, completed, remaining, nextCursor,
usedSerializedBytes, and maxSerializedBytes; results retain requestIndex and isolate
individual errors. The complete UTF-8 JSON response, including its envelope and cursor, stays
within the requested budget and the reported 49,152-byte ceiling.
The first call pins every resolved revision before canonical bodies load; cursor calls keep
those pins, so concurrent writes cannot mix revisions. Its response lists every requested
requestIndex, while a cursor response is sparse and lists only the indexes it touched; an
omitted index is neither a failure nor a completion. Admit whole compact messages in
deterministic round-robin order and loop on the returned cursor until nextCursor is null:
Per-item continuation values remain available for compatibility, but are not the primary
workflow. An oversized message returns bounded page.oversizedMessage metadata
(conversationId, revisionId, sourceNodeId, offset, bytes) without text and advances
cursor state; recover its complete content through an authorized canonical HTTP read or account
export rather than retrying the same batch page. Single reads accept format: "compact";
canonical output remains the default. memory_resolve_conversations resolves up to 20 known
conversation owners by exact title without semantic search, returning conversation IDs, current
revision IDs, and live tags. memory_list_revisions returns the immutable revision
history of one owned conversation (metadata only, newest first, cursor-paged), so a client can
pin and read any earlier revision with memory_get_conversation instead of relying on a
retained write receipt. memory_restore_revision restores an owned conversation to any historical
revision with base_revision_id optimistic concurrency, recording an immutable head transition
without creating redundant canonical revisions. memory_copy_conversations performs a lossless
canonical R2 copy of 1–20 owned conversations into another owned namespace using a required
idempotency_key and attaching first-class derivedFrom provenance, rather than a compact-message
restorable via memory_store. Store/append/replace/restore/copy accept verify: true to reload the
committed R2 revision and return compact readback with separate indexing/verification status.
Store, append, replace, and restore return a bounded durable receipt: a committed revision reports
durable: true with separate indexing and verification status, and the whole receipt is fitted to a
documented safe maximum of 49,152 bytes (48 KiB) below the 64 KiB MCP tool guard — the deployed
max_receipt_bytes and max_tool_output_bytes values from memory_get_capabilities. Inline readback is
returned only when it fits; otherwise the receipt carries readback_requests — selectors that are
directly reusable as memory_get_conversations first-call requests (loop nextCursor until null)
— plus an omitted list naming the shed field paths. Fitting sheds verbose fields in a fixed order and
never drops identity, durable, or status fields, so a durable committed mutation never becomes a
generic response-size error.
See the RP workflow and reviewable runtime-rule amendment.
For prompt and task execution, memory_build_context compiles a deterministic, revision-pinned context pack
from 1–20 required canonical conversations (exact title or conversation ID, full or tail mode,
active or all branch, and optional follow pointer-expansion fields) and up to 8 optional hybrid-search retrieval
requests within caller-specified token and serialized-byte limits (at most max_response_bytes,
49,152 bytes / 48 KiB, from memory_get_capabilities). It pins required
current revisions before loading canonical bodies, follows exact structured pointers (such as current_scene and
active_arc.owner) deterministically without semantic fallback, deduplicates source messages structurally on
(conversation_id, revision_id, source_node_id), orders content deterministically by authority (explicit required,
then required expanded, then optional expanded, then retrieved), priority, score, and stable IDs, and greedily fits
whole messages. If required content alone exceeds either budget, it returns a bounded diagnostic with suggested minimums
without leaking canonical text. Context packs are extractive-only, include full message provenance and an optional
deterministic compiled text projection, have a deterministic pack_id, and perform no writes or generative summarization.
This runs formatting, lint, strict TypeScript (Worker and the web/mindmap/ browser project), the memory-map bundle freshness check, unit/MCP/retrieval tests, Workers-runtime D1/R2 integration tests, and a Wrangler deploy dry run. No command deploys unless yarn deploy is invoked explicitly.
After changing anything under web/mindmap/, run yarn build:mindmap so the generated src/mindmap-bundle.ts matches its sources.
Reindexing reads canonical R2 data; the ChatGPT export does not need to be uploaded again. D1 migrations are numbered SQL files and must be applied through Wrangler—never by dashboard drift.