# MemPersist

**Category:** 🧠 Knowledge & Memory  
**Repository:** https://github.com/ravhirizaldi/mempersist  
**Views:** 0  
**Installs:** 0  
**Upvotes:** 0  
**Directory Page:** https://allmcps.com/mcp/mempersist

## Description
Durable, revision-pinned memory storage and retrieval for AI conversations.

## Claude Desktop Quick Installation
Heuristic fallback — verify the package name and runner against the repository README before running it. Uses `npx` (confidence: low):

```json
"mcpServers": {
  "mempersist": {
    "command": "npx",
    "args": ["-y","mempersist"]
  }
}
```

## Documentation & README

# Mempersist

Mempersist is a clean-room, Cloudflare-native long-term memory service for AI conversations. It preserves original ChatGPT exports and normalized conversation graphs in R2, catalogs them in D1, builds disposable lexical and semantic indexes, and exposes compact retrieval and intentional writes through MCP.

It does not silently capture ChatGPT traffic, extract replacement “facts,” provide a SaaS billing layer, or make search indexes canonical. “Unlimited” means no application message quota; Cloudflare limits and billing still apply.

## Architecture

```text
ChatGPT conversations.json / MCP writes
                  |
          validation + IDs
                  |
          R2 canonical archive  <------ export / recovery
                  |
             D1 catalog
                  |
        Cloudflare Queues
          /             \
   D1 FTS5          Workers AI BGE-M3 -> Vectorize
          \             /
       normalized hybrid search
                  |
       HTTP + OAuth-protected MCP
```

R2 is the source of truth. D1 holds operational metadata and the derived FTS representation. Vectorize is disposable. A canonical write succeeds before indexing is queued, and an indexing failure never reports that durable memory was lost.

See [ARCHITECTURE.md](https://github.com/ravhirizaldi/mempersist/blob/HEAD/ARCHITECTURE.md), [SECURITY.md](https://github.com/ravhirizaldi/mempersist/blob/HEAD/SECURITY.md), and [docs/operations-and-recovery.md](https://github.com/ravhirizaldi/mempersist/blob/HEAD/docs/operations-and-recovery.md).

## Prerequisites

- WSL2/Linux, Node.js 22+, Yarn 1.22, and Wrangler 4.x
- A Cloudflare account with Workers, D1, R2, Vectorize, Workers AI, and Queues available
- Wrangler OAuth authentication: `yarn wrangler whoami`

Use Yarn only.

## Setup

```bash
yarn install
cp .dev.vars.example .dev.vars
yarn types:bindings
yarn db:migrate:local
yarn dev
```

Set a long random `MEMORY_API_TOKEN` in `.dev.vars`. Local D1 and R2 are simulated; Workers AI and Vectorize bindings are remote in the main configuration. Unit and integration tests do not call remote AI.

## Cloudflare provisioning

Provisioning is intentionally manual and must be explicitly authorized. Follow [docs/cloudflare-resources.md](https://github.com/ravhirizaldi/mempersist/blob/HEAD/docs/cloudflare-resources.md), then add the real D1 `database_id` returned by Wrangler to `wrangler.jsonc`. Never invent IDs or reuse unrelated account resources.

Set the production secret without putting it in source:

```bash
yarn wrangler secret put MEMORY_API_TOKEN
```

Apply migrations and deploy only after review:

```bash
yarn db:migrate:remote
yarn deploy:dry-run
yarn deploy
```

## ChatGPT import

Export data from ChatGPT, extract `conversations.json`, then:

```bash
MEMPERSIST_URL=http://localhost:8787 \
MEMPERSIST_TOKEN='your-token' \
yarn import:chatgpt /path/to/conversations.json
```

Files up to 16 MiB use direct streaming upload. Larger files use 16 MiB R2 multipart parts — the
deployed `max_direct_import_bytes` and `max_multipart_part_bytes` values from
`memory_get_capabilities`. The Worker hashes the completed object, preserves it unchanged, detects
exact duplicate exports, and processes at most 25 conversations per queue turn. Check progress with:

```bash
MEMPERSIST_TOKEN='your-token' yarn import:status <import-id>
```

See [docs/chatgpt-import.md](https://github.com/ravhirizaldi/mempersist/blob/HEAD/docs/chatgpt-import.md).

## MCP

The Streamable HTTP endpoint is `https://<worker>/mcp`. Interactive clients such as ChatGPT
use OAuth 2.1 authorization-code flow with PKCE: the consent page takes an email, sends a
single-use magic link, and completes the connection only after the link is opened. An existing
email reconnects to its archive; a new archive is created after the first link. Each account can
own multiple namespaces, and the same namespace name may exist in different accounts with fully
separated data.
Developer scripts and the CLI may keep sending `MEMORY_API_TOKEN` as a bearer token for the
owner archive.

Email authentication uses the `EMAIL` send binding. The sending address per endpoint is
configuration (`AUTH_EMAIL_FROM`, `LEGACY_AUTH_EMAIL_FROM`), not part of the API contract.

The public site, OAuth pages, and magic-link email support English and Bahasa Indonesia. Use the
language switcher to persist a browser preference; otherwise MemPersist uses `Accept-Language` and
falls back to English. API, MCP, and CLI contracts remain English.

## Dashboard

Open `/login` for passwordless access to the server-rendered dashboard. It includes archive totals,
a deterministic memory map, revision-pinned canonical conversation reading, display-name editing,
and a streamed lossless export of current memory. Namespace emptying runs asynchronously. Account
deletion has a seven-day cancelable grace period and makes writes read-only while pending.

See [docs/dashboard.md](https://github.com/ravhirizaldi/mempersist/blob/HEAD/docs/dashboard.md) and ADR 0028. No frontend framework, extra dependency,
or additional Cloudflare resource is required.

For the deployed Worker, add `https://mempersist.codifiedtech.id/mcp` as a custom MCP app in
ChatGPT Developer mode. ChatGPT discovers OAuth automatically, opens the consent page, and
stores the issued access/refresh tokens. Existing connections keep working after upgrades
without re-authorization. Clients already configured with the legacy endpoint remain
supported; changing one to the primary endpoint requires one new authorization. Do not paste
`MEMORY_API_TOKEN` into ChatGPT's connector settings.

Available tools:

- `memory_search`
- `memory_get_context`
- `memory_get_conversation`
- `memory_get_conversations`
- `memory_list_conversations`
- `memory_list_revisions`
- `memory_resolve_conversations`
- `memory_build_context`
- `memory_list_namespaces`
- `memory_stats`
- `memory_store`
- `memory_append`
- `memory_replace`
- `memory_restore_revision`
- `memory_copy_conversations`
- `memory_get_capabilities`

`memory_store` and `memory_append` accept optional tags (lowercased, deduplicated, up to 20);
`memory_search` filters by tags with AND semantics and returns each conversation's tags. See
[docs/mcp.md](https://github.com/ravhirizaldi/mempersist/blob/HEAD/docs/mcp.md) and ADR 0013.

- `memory_delete_conversations`
- `memory_empty_namespace`
- `memory_import_status`

Search returns compact references; call `memory_get_context` only for selected results. See [docs/mcp.md](https://github.com/ravhirizaldi/mempersist/blob/HEAD/docs/mcp.md).

`memory_get_capabilities` returns the deployed capability contract: protocol and capabilities
versions, transport and message limits, per-tool item and byte bounds, and feature flags. The
limits quoted below are the same enforced values it reports, so treat that tool as the single
contract rather than independent figures. Aggregate limits are measured as UTF-8 bytes of the
complete serialized request before any canonical write begins: inline JSON writes on both
transports are limited to 1,048,576 bytes (1 MiB), and imports are limited to 16 MiB direct and
16 MiB per multipart part. An oversized request fails with code `REQUEST_TOO_LARGE` and a
conservative `suggested_max_items` estimate, never a silent truncation.

For known memories, start `memory_get_conversations` with a `requests` array of 1–20
conversation requests and optional `max_serialized_bytes` (default 32,768; minimum 4,096;
maximum 49,152 — the deployed values `memory_get_capabilities` reports). A continuation sends
exactly one opaque `cursor` plus the optional budget —
never both `requests` and `cursor`, and never neither. Results use the established camelCase
batch fields `batchId`, `results`, `completed`, `remaining`, `nextCursor`,
`usedSerializedBytes`, and `maxSerializedBytes`; `results` retain `requestIndex` and isolate
individual errors. The complete UTF-8 JSON response, including its envelope and cursor, stays
within the requested budget and the reported 49,152-byte ceiling.

The first call pins every resolved revision before canonical bodies load; cursor calls keep
those pins, so concurrent writes cannot mix revisions. Its response lists every requested
`requestIndex`, while a cursor response is sparse and lists only the indexes it touched; an
omitted index is neither a failure nor a completion. Admit whole compact messages in
deterministic round-robin order and loop on the returned cursor until `nextCursor` is null:

```text
call memory_get_conversations({ cursor: nextCursor })
```

Per-item `continuation` values remain available for compatibility, but are not the primary
workflow. An oversized message returns bounded `page.oversizedMessage` metadata
(`conversationId`, `revisionId`, `sourceNodeId`, `offset`, `bytes`) without text and advances
cursor state; recover its complete content through an authorized canonical HTTP read or account
export rather than retrying the same batch page. Single reads accept `format: "compact"`;
canonical output remains the default. `memory_resolve_conversations` resolves up to 20 known
conversation owners by exact title without semantic search, returning conversation IDs, current
revision IDs, and live tags. `memory_list_revisions` returns the immutable revision
history of one owned conversation (metadata only, newest first, cursor-paged), so a client can
pin and read any earlier revision with `memory_get_conversation` instead of relying on a
retained write receipt. `memory_restore_revision` restores an owned conversation to any historical
revision with `base_revision_id` optimistic concurrency, recording an immutable head transition
without creating redundant canonical revisions. `memory_copy_conversations` performs a lossless
canonical R2 copy of 1–20 owned conversations into another owned namespace using a required
`idempotency_key` and attaching first-class `derivedFrom` provenance, rather than a compact-message
restorable via `memory_store`. Store/append/replace/restore/copy accept `verify: true` to reload the
committed R2 revision and return compact readback with separate indexing/verification status.

Store, append, replace, and restore return a bounded durable receipt: a committed revision reports
`durable: true` with separate indexing and verification status, and the whole receipt is fitted to a
documented safe maximum of 49,152 bytes (48 KiB) below the 64 KiB MCP tool guard — the deployed
`max_receipt_bytes` and `max_tool_output_bytes` values from `memory_get_capabilities`. Inline readback is
returned only when it fits; otherwise the receipt carries `readback_requests` — selectors that are
directly reusable as `memory_get_conversations` first-call `requests` (loop `nextCursor` until `null`)
— plus an `omitted` list naming the shed field paths. Fitting sheds verbose fields in a fixed order and
never drops identity, durable, or status fields, so a durable committed mutation never becomes a
generic response-size error.
See the [RP workflow and reviewable runtime-rule amendment](https://github.com/ravhirizaldi/mempersist/blob/HEAD/docs/rp-workflow.md).

For prompt and task execution, `memory_build_context` compiles a deterministic, revision-pinned context pack
from 1–20 required canonical conversations (exact title or conversation ID, full or tail mode,
active or all branch, and optional `follow` pointer-expansion fields) and up to 8 optional hybrid-search retrieval
requests within caller-specified token and serialized-byte limits (at most `max_response_bytes`,
49,152 bytes / 48 KiB, from `memory_get_capabilities`). It pins required
current revisions before loading canonical bodies, follows exact structured pointers (such as `current_scene` and
`active_arc.owner`) deterministically without semantic fallback, deduplicates source messages structurally on
`(conversation_id, revision_id, source_node_id)`, orders content deterministically by authority (explicit required,
then required expanded, then optional expanded, then retrieved), priority, score, and stable IDs, and greedily fits
whole messages. If required content alone exceeds either budget, it returns a bounded diagnostic with suggested minimums
without leaking canonical text. Context packs are extractive-only, include full message provenance and an optional
deterministic compiled text projection, have a deterministic `pack_id`, and perform no writes or generative summarization.

## Quality gate

```bash
yarn verify
```

This runs formatting, lint, strict TypeScript (Worker and the `web/mindmap/` browser project), the memory-map bundle freshness check, unit/MCP/retrieval tests, Workers-runtime D1/R2 integration tests, and a Wrangler deploy dry run. No command deploys unless `yarn deploy` is invoked explicitly.

After changing anything under `web/mindmap/`, run `yarn build:mindmap` so the generated `src/mindmap-bundle.ts` matches its sources.

## Operations

```bash
yarn admin search 'api.internal.example'
yarn retry <job-id>
yarn reindex
yarn verify:integrity
```

Reindexing reads canonical R2 data; the ChatGPT export does not need to be uploaded again. D1 migrations are numbered SQL files and must be applied through Wrangler—never by dashboard drift.

