The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Polymnemo listing page.
A shared long-term memory across any LLM, over MCP.
Point Claude Desktop, an MCP-capable IDE, or any MCP client at one polymnemo endpoint and they share the same memories — stored in your own Postgres. Store a fact with one assistant, recall it from another; save whole sessions and reload them; even attach files, images, or video. Embeddings run locally (no embedding API key), and the server makes no generative-LLM calls.
Status: active development. Semantic memory + pgvector store, session save/reload, and multimedia memories all work; deployable to Azure Container Apps (or Cloud Run) via Terraform. Wiki · Issues
Protocols.A request carries a bearer key (which resolves to a user_id); the tool passes
the rate-limit gate, then delegates to a MemoryService that chunks + embeds
text and stores the vectors in pgvector — large files go to object storage via
presigned URLs, with only a searchable description embedded.
Every layer is a typing.Protocol, wired together by a composition root
(context.py), so you can swap an implementation
without touching the tools:
| Layer | Default | Swap for |
|---|---|---|
| Store | PostgresStore (pgvector) | InMemoryStore (dev/tests) |
| Embedder | fastembed (local ONNX) | StubEmbedder (offline) |
| Auth | GitHub OAuth + API tokens | static bearer keys (dev) |
| Retriever | VectorRetriever | your own ranker |
| BlobStore | S3 / R2 | off |
| RateLimiter | global token bucket | off |
The ping tool returns the active layers, so you can see how a running server is
wired.
Requires Python 3.11+.
The durable store is Postgres + pgvector; the easiest hosted option is Neon (use the pooled connection string). Apply the schema once:
Copy .env.example to .env and set the database URL and at least one API key:
Each key maps a bearer token to a user_id; writes are scoped to that user.
Or with Docker:
Any client that supports remote (HTTP) MCP servers with custom headers needs two things:
http://<host>:<port>/mcpAuthorization: Bearer <your-key>For clients that read an mcpServers config:
Or verify with the inspector:
user_id, and writes are owner-scoped (you can only
edit or delete your own memories).<server>/tokens to create, list and revoke them; a new token is shown
once and stored only as a SHA-256 hash. Deliberately not MCP tools — a
secret must never travel an LLM channel.shared) is readable by everyone ("born shared"); everything else is private
to its owner. Sessions and media default to private namespaces.remember may return several ids and recall returns the closest chunks.polymnemo exposes MCP tools for storing, searching, and managing memories:
remember, recall, list_memories, get_memory, update, forgetsave_session, load_sessioncreate_upload, confirm_upload, get_download_urlPlus a ping health check and a memory://{namespace} resource for
auto-injecting a collection. Media bytes go to object storage via presigned URLs
— never through the MCP channel — with only a searchable description embedded
(needs the blob extra).
Full arguments and return shapes live in the dedicated MCP tools reference (coming soon). See Concepts for how keys, namespaces, and chunking work.
POLYMNEMO_* environment variables (or .env) — full list in
.env.example. The essentials:
| Variable | Default | Purpose |
|---|---|---|
POLYMNEMO_DATABASE_URL | (unset) | Postgres+pgvector DSN. Required in production; unset → in-memory (dev/tests). |
POLYMNEMO_API_KEYS | (empty) | "key1:alice,key2:bob" — required for bearer auth. |
POLYMNEMO_SHARED_NAMESPACES | shared | Namespaces readable by every user. |
POLYMNEMO_HOST / POLYMNEMO_PORT / POLYMNEMO_MCP_PATH | 127.0.0.1 / 8000 / /mcp | Transport. |
POLYMNEMO_RATELIMIT_ENABLED / _PER_MIN | false / 600 | Optional global rate limit (ops per minute). |
POLYMNEMO_BLOB_BACKEND (+ _BUCKET / _ENDPOINT_URL / _ACCESS_KEY_ID / _SECRET_ACCESS_KEY) | none | Object storage for media memories; s3 = Cloudflare R2 / S3-compatible. |
Postgres tests run when TEST_DATABASE_URL points at a pgvector Postgres. CI
(.github/workflows/ci.yml) runs lint + the suite with coverage and posts a
pass/fail/coverage table to each run's summary.
polymnemo is stateless (all state in Neon + object storage), so it runs on
Azure Container Apps (primary) or Google Cloud Run with scale-to-zero.
Everything is Terraform in deploy/terraform/: shared
neon/ (Postgres) and r2/
(media bucket) roots own the durable state, and a compute root deploys a service
that reads both — so memories and media are shared across clouds.
The Azure path deploys from CI in one click: set the GitHub secrets
(scripts/setup-github-secrets.sh), bootstrap the state backend
(scripts/bootstrap-tfstate-azure.sh), then run the deploy (azure) workflow
(neon → schema → r2 → build → app). GCP is a manual failover. Full walkthrough
in docs/deploy.md.
MIT — see LICENSE.