The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the DeepMem listing page.
Drop-in AI memory layer with 2× faster response and 10× lower cost.
Fully compatible with Mem0 API. Migrate in 5 minutes - one import line.
Self-hostable. No auth, no payment, no lock-in. Or use the managed cloud at deepmem.dev.
Cloud · Self-host · Benchmarks · Reproduce them
Migrate from Mem0 in one line - same MemoryClient, same method signatures:
Turn conversations into searchable long-term memory: a FastAPI HTTP API in
front of a Qdrant vector store, with LLM fact extraction, hybrid retrieval
(vector + BM25 + entity boost + time-decay), semantic caching, async batched
distillation, GDPR controls, and a built-in MCP server. It runs in open
mode - no API key, no user registration - so you can deploy it for your own
agents in minutes. Multi-tenant isolation is driven by user_id in the
request body.
Prefer not to self-host? DeepMem Cloud is the managed version of this exact engine at deepmem.dev - same API, no infra. Sign up, grab a key (
dm_live_...), point your base URL athttps://deepmem.dev, done. The cloud and the open-source server speak the same Mem0-compatible API, so client code is identical.
Already using Mem0? Switch to DeepMem cloud in one line. The deepmem-client
package mirrors mem0.MemoryClient - same class name, same method signatures,
same filters={"user_id": ...} style - so everything after the import stays
untouched.
add(infer=True) (the default) is asynchronous on DeepMem cloud - it
returns pending=True with results=[] and extracted facts land a few
seconds later. (Mem0 cloud's add is async too - it returns PENDING.) Pass
infer=False for synchronous raw-text storage that's immediately searchable.relations is always []. Mem0's graph features aren't
replicated.reset differs - Mem0's is account-wide; DeepMem's is per-user_id with
a confirm guard.DeepMem Cloud is 10x cheaper than Mem0 at every paid tier - the same shape of plans, a tenth of the price.
| Tier | DeepMem | Mem0 cloud |
|---|---|---|
| Hobby | Free | Free |
| Starter | $1.9/mo | $19/mo |
| Growth | $7.9/mo | $79/mo |
| Professional | $24.9/mo | $249/mo |
Self-host instead and it's $0 - you pay only your own LLM/embedding provider (the same LLM cost Mem0 charges on top of its plan price), with no memory-service markup. Batched distillation also cuts LLM calls ~80%, so even your provider bill is smaller than per-message extractors.
Plans and limits: deepmem.dev · mem0.ai.
No cherry-picked headline. The scripts and workload ship in
/benchmarks - run them yourself. Here's what we
measured and the exact config that produced it:
| Metric | DeepMem self-hosted ¹ | DeepMem cloud | Mem0 cloud |
|---|---|---|---|
| Search p50 | 73 ms | 643 ms | 653 ms |
| Search p95 | 86 ms | 811 ms | 710 ms |
| Search hits (40 queries) | - | 195 | 84 |
| Add p50 (raw store) | 899 ms ² | 792 ms | 695 ms ³ |
¹ BGE-M3 on a GTX 1070 GPU (2016-era), local file Qdrant,
infer=False, 100 ops, concurrency 1. ² Dominated by local-file Qdrant I/O - a Qdrant server cuts this sharply. ³ Mem0 has no raw-store mode;addalways runs LLM extraction, so this row isn't apples-to-apples.
/benchmarks from a low-RTT location for your own
numbers.Agent frameworks keep re-discovering that they need persistent, retrievable memory. The hosted options bill per call and send your data to someone else's cloud. DeepMem is the self-hostable alternative: the same Mem0-shaped API you can drop in, but it runs on your box, with your embedder, your LLM key, and your Qdrant - and the code is right here to verify it.
How does DeepMem compare to other Mem0 alternatives? Most are hosted-only or layer memory on top of someone else's vector DB. DeepMem combines three things at once: it's self-hostable (your data stays on your box - $0 beyond your own LLM key), MCP-native (Claude Desktop / Cursor read and write memories directly as tools), and fully open-source - and the managed cloud runs the exact same engine, so cloud and self-host are one API, not two products.
| Without DeepMem | With DeepMem |
|---|---|
| Re-explain who you are and what you're working on every session | The agent recalls identity, projects, and preferences automatically |
| Lose debugging and research context between sessions | Past root causes, dead ends, and findings are recalled, so work isn't repeated |
| Manually restate preferences every session | Preferences persist across sessions, agents, and projects |
| Hosted memory services that bill per call and hold your data | Self-host on your infra, or use the cloud - your call, same API |
Three ways to run. All speak the same Mem0-compatible API.
Or pull the published image:
The image exposes :8000 (HTTP) and :8001 (MCP). The Dockerfile
and docker-compose.yml cover the GPU variant (CUDA torch + BGE_DEVICE=cuda)
and BGE-M3 model-download options (HF mirror, proxy, or local mount).
Write and search in three lines:
user_idis optional (defaults to"default"); send differentuser_ids to isolate end-users.infer: falsestores raw text immediately (test-friendly); the defaultinfer: truequeues for LLM fact extraction.
Multi-provider by config, not code. Both layers switch on env vars:
| Layer | Options | Selector |
|---|---|---|
| LLM (fact extraction) | OpenAI · Anthropic (native SDK) · any OpenAI-compatible (DeepSeek / vLLM / Ollama / Groq / LM Studio) | LLM_PROVIDER + LLM_API_KEY / ANTHROPIC_API_KEY / DEEPSEEK_API_KEY |
| Embeddings | BGE-M3 (local, GPU/CPU) · Google Gemini · any OpenAI-compatible | EMBEDDING_PROVIDER + BGE_M3_PATH / GOOGLE_API_KEY / OPENAI_API_KEY |
BGE_DEVICE=auto|cpu|cuda picks GPU when available, else CPU (force cpu
on small-VRAM cards to avoid multi-process contention). BYOK overrides the
LLM per-request.
Hybrid retrieval. Vector similarity + BM25 keyword + entity boost + time-decay, fused into one score. Over-fetch, re-rank, return.
Async batched distillation. Writes queue behind a silence window and are extracted in batches - ~80% fewer LLM calls than per-message extraction.
Semantic cache. Repeat adds/searches hit a similarity-gated cache and return cached facts without re-embedding or re-querying Qdrant.
Stores preferences, not code. extraction_filter strips large fenced
code blocks before LLM extraction, so the store fills with durable facts,
not pasted implementations.
MCP server. deepmem_write / deepmem_search / deepmem_delete tools
for Claude Desktop, Cursor, and any MCP client.
GDPR. Soft-delete with retention window, hard-delete reset, SHA-256
id masking in logs, export/import for portability.
| Method | Path | Description |
|---|---|---|
| POST | /v1/memories | Write messages; LLM-extract facts (infer=false stores raw) |
| POST | /v1/memories/search | Semantic search (vector + BM25 + entity + time-decay) |
| GET | /v1/memories | List all for a user_id (paginated) |
| GET | /v1/memories/{id} | Get one by ID |
| PUT | /v1/memories/{id} | Update one memory's text |
| DELETE | /v1/memories/{id} | Soft-delete one |
| DELETE | /v1/memories | Soft-delete all for a user_id (GDPR) |
| GET | /v1/memories/{id}/history | ADD/UPDATE/DELETE audit log |
| POST | /v1/reset | Hard-delete all + history (needs confirm_user_id) |
| GET | /v1/export · POST /v1/import | Portable JSON export / import |
| GET | /health · /ready | Liveness / readiness probes |
agent_id / run_id optionally scope writes/reads (mirrors Mem0's three-level
isolation: user -> agent -> run). Interactive docs at /docs.
The cloud exposes a remote MCP endpoint (streamable HTTP):
Works with Claude Code, Claude Desktop, Cursor, and any MCP-compatible client.
Claude Code - one command:
Or install the plugin (bundles a /deepmem:setup command that guides API-key
configuration):
Cursor → Settings → MCP → Add new MCP server, or edit
~/.cursor/mcp.json (global) / .cursor/mcp.json (per project). Cursor
resolves ${env:NAME} variables in url and headers, so the key can live
in an environment variable instead of the config file:
(with DEEPMEM_API_KEY=dm_live_... exported in your shell, or paste the
literal key if you prefer). Community MCP listings with one-click "Add to
Cursor" live on cursor.directory; official
plugins ship in the Cursor Marketplace.
DeepSeek Harness (dsh) - one overlay file, no dsh plugin needed (dsh
connects any MCP server through its generic dsh-mcp-client bridge):
See examples/dsh for details and self-hosted variants.
Self-hosted (stdio, bundled server, no key needed in open mode):
This repo ships a .cursor/mcp.json preconfigured for the self-hosted server - open the repo in Cursor and enable it under Settings → MCP.
Single Qdrant collection (memories), hard-filtered by user_id payload.
TenantValidator NFC-normalizes and enforces a [A-Za-z0-9._:-]{1,256}
charset on user_id - never trust the raw request value.
Config loads once from env vars (.env, auto-loaded) > config.json > defaults.
Key variables:
| Variable | Required | Description |
|---|---|---|
DEEPSEEK_API_KEY / LLM_API_KEY / ANTHROPIC_API_KEY | one LLM key | LLM for fact extraction |
LLM_PROVIDER | no | auto / openai / anthropic / openai_compatible (default auto) |
EMBEDDING_PROVIDER | no | bge-m3 / google / openai (default bge-m3) |
BGE_M3_PATH | no | local BGE-M3 dir or HF model id (default BAAI/bge-m3) |
BGE_DEVICE | no | auto / cpu / cuda (default auto) |
QDRANT_URL / QDRANT_API_KEY | no | remote Qdrant; omit for local file store |
CORS_ORIGINS | yes | comma-separated allowed origins (no * in prod) |
RATE_LIMIT_ADD / RATE_LIMIT_SEARCH | no | per-minute limits (default 30 / 60) |
Backends auto-switch on env vars - no code changes:
QDRANT_URL set -> remote Qdrant; unset -> local file Qdrant under ./data/qdrant.CORS_ORIGINS=* is refused at boot unless DEEPMEMORY_DEBUG=1.For HTTPS, put Caddy or Nginx in front; scripts/start.sh runs under systemd
or any process manager.
Two reproducible scripts in /benchmarks:
python benchmarks/benchmark_cloud.py.python benchmarks/run_benchmark.py.Both ship a self-contained workload and report P50/P95/P99 + throughput. The numbers at the top of this README were produced with these scripts - rerun them and read your own percentiles.
Do I need the cloud? No. The open-source server is fully functional on its own. The cloud (deepmem.dev) is the zero-ops option - same API.
Does it work offline? Retrieval and raw-store (infer=false) work with no
network. LLM fact extraction (infer=true) needs an LLM key - or run a local
OpenAI-compatible model (Ollama / vLLM / LM Studio) and point LLM_BASE_URL
at it.
Where is my data? In your Qdrant (local file or server) + a SQLite audit log. Nothing leaves your machine except the LLM/embedding calls you configure.
Multi-tenant? Yes - single Qdrant collection hard-filtered by user_id.
agent_id / run_id add agent and session scope.
Is it production-ready? Used in production under systemd with a remote Qdrant. Local file Qdrant is fine for dev/single-worker; use a Qdrant server for multi-worker or high-throughput.
MIT.