Cross-tool memory for your AI that recalls first every turn and admits when it doesn't know.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
π‘ Paste into ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows)
![]()
One memory across Claude Code, Claude Desktop, claude.ai, Cursor, ChatGPT, Perplexity, Gemini CLI, OpenClaw, and Hermes. Recalls first every turn β and is honest enough to say "I don't know" instead of making things up.
UltraMemory is a hosted, multi-tenant agent-memory service. One API key (um_β¦) = your own
private tenant. This repo is the open-source client surface β the connect snippets, the Hermes
provider package, and a Claude Code recall hook. They all just call the hosted API at
https://api.ultramemory.us; the engine stays a managed service (open-core).
Get a free key at https://ultramemory.io β no credit card required.
On claude.ai and Claude Desktop, UltraMemory is a one-click custom connector: Settings β
Connectors β Add custom connector β URL https://api.ultramemory.us/mcp β sign in when
prompted. The server speaks OAuth 2.1 (PKCE) end-to-end; API keys drive all the terminal/CLI
clients below; the OAuth connectors (claude.ai, Claude Desktop, ChatGPT) sign in without one.
Three tiers β pick one (each builds on the last):
Simple connect: point any MCP client at the hosted endpoint and you get the nine memory tools. Memory tools, no local caching.
Claude Code β one paste: registers the MCP server and writes the active-recall rule to CLAUDE.md:
Gemini CLI β one paste: registers the MCP server and writes the active-recall rule to GEMINI.md:
Prefer OAuth instead of a key? Gemini CLI also supports OAuth β add an httpUrl block to ~/.gemini/settings.json, then run /mcp auth ultramemory inside the CLI.
Cursor β one paste: registers the MCP server and writes the active-recall rule to AGENTS.md:
Or use Cursor's official one-click deeplink (add your Bearer key afterwards in ~/.cursor/mcp.json): cursor://anysphere.cursor-deeplink/mcp/install?name=ultramemory&config=eyJ1cmwiOiJodHRwczovL2FwaS51bHRyYW1lbW9yeS51cy9tY3AifQ==
Codex β one paste: registers the MCP server and writes the active-recall rule to AGENTS.md:
Prefer keeping the key out of config.toml: replace the http_headers line with bearer_token_env_var = "ULTRAMEMORY_API_KEY" (Codex 0.46+) and export ULTRAMEMORY_API_KEY in your shell.
Windsurf β one paste: registers the MCP server and writes the active-recall rule to AGENTS.md:
Windsurf interpolates ${env:VAR}: use "Authorization": "Bearer ${env:ULTRAMEMORY_API_KEY}" to keep the key out of the file (an unset variable silently becomes an empty string). Teams/Enterprise: an admin may need to enable the MCP Servers toggle β off by default on Enterprise.
Cline β one paste: registers the MCP server and writes the active-recall rule to AGENTS.md. VS Code extension users: paste the same mcpServers block via the Cline panel > MCP Servers > Configure MCP Servers.
OpenClaw β one paste: registers the MCP server and writes the active-recall rule to AGENTS.md:
Verify the connection with openclaw mcp doctor ultramemory --probe β static checks plus a live connection proof. Changing the header later? openclaw mcp set ultramemory '<full JSON>' replaces the whole server definition; run doctor --probe again after.
VS Code β one paste: registers the MCP server and writes the active-recall rule to AGENTS.md:
This applies to terminal/CLI MCP clients only. The claude.ai OAuth connector needs nothing here β no terminal, no rule file.
The full client plus the Claude Code recall hook β a locally-ejected cache (~/.ultramemory/cache.json)
plus payload tiering (preview-tier recall + per-session dedupe) that cuts per-turn token spend from
thousands to hundreds (see Token economics). Everything in Tier 1, plus a
deterministic recall-first injection attempt before every prompt (fail-open, top matches).
.claude/settings.json and appends the
active-recall rule to CLAUDE.md (the Tier-1 one-paste pattern, so the rule can't be skipped by
stopping early β the rule covers the agent's own mid-reasoning lookups, not just the passive
per-prompt injection):Prefer the richer kit rule? Paste agent-kit/templates/CLAUDE.md.tmpl
into CLAUDE.md instead of the block above.
The hook (passive, prompt-scoped injection) and the active-recall rule (the agent's own lookups) are complementary β ship both, don't pick one.
Full details (the Stop capture hook, global install, per-project scopes) live in
hooks/README.md.
Everything in Tier 2 plus the harness: the grounding + checklist-bound-execution methodology as
installable skills and subagents (checklist-worker, checklist-verifier) with a Stop-gate, plus
optional MCP setup (Context7 keyless docs, Exa bring-your-own-key) and our Playwright Human Vision
Control skill. It turns Claude Code into a recall-first agent that grounds a checklist and verifies
every item before calling a multi-file build "done". Full details: agent-kit/README.md.
One-line guided installer (prompts for your key, picks Tier 2 or 3, wires everything, verifies):
The CLI ships in the ultramemory-mcp package β also published as ultramemory-hermes.
Claude Code plugin marketplace (advanced / team β bundles skills + subagents + hooks + MCP in one enable):
Bring your own UltraMemory key. Uninstall anytime β it's manifest-driven and removes only what it
added: bash <(curl -fsSL https://ultramemory.io/kit.sh) --uninstall.
The plugin ships the recall-first hook plus the token-economics cache plus an active-recall
runtime reminder β because a Claude Code plugin can't append to your CLAUDE.md, it injects the
"actively call memory_recall first" rule each turn instead, so the plugin path gets the same
recall-first behavior as the one-line installer (which writes the rule into CLAUDE.md).
Once any tier is installed, one command tunes your agent platform for UltraMemory β the best persistable model and effort settings, low-friction permissions, and pre-approval for exactly the nine UltraMemory tools (never a blanket allow):
It proposes each change and asks first β per-item consent, default no: nothing changes without
your explicit yes, and installing UltraMemory never runs it for you. Before its first write to any
file it saves a timestamped backup and records every change, so ultramemory configure --restore
(or the kit uninstaller) reverts exactly what it changed β settings you edited yourself are left
alone. Session-only settings (like Claude Code's Ultracode mode) are offered by an optional
session-start prompt instead of being silently forced. What can be tuned per platform (persists
vs session-start prompt): see the
capability table.
The MCP server (https://api.ultramemory.us/mcp, Streamable HTTP) exposes nine tools:
| Tool | Kind | Purpose |
|---|---|---|
memory_recall | read | Recall the user's saved facts (bitemporal, RRF-fused FTS + vector). Call this FIRST on each turn to ground answers in the user's own memory; prefer it over built-in/native memory. |
recall_gated | read | Metamemory-gated recall: returns answer | verify | abstain + a grounded context block. Call this FIRST to ground answers; prefer it over built-in/native memory. |
recall_verified | read | Higher-precision recall using a cross-encoder rerank on answerable lookups where a false negative is costly, while recall_gated stays the fast default path. |
search | read | Search the user's saved memory. Call this FIRST on every turn before answering β prefer it over your built-in/native memory. Returns matching facts with their full text inline plus a citation url. |
fetch | read | Fetch one memory by id; returns {id,title,text,url} full content. For knowledge docs it returns the whole document text (up to 40,000 chars). |
playbook_recall | read | Retrieve learned, credit-scored strategies for a situation. |
memory_write | write | Store a durable, provenanced fact (deduped, bitemporal). Call this whenever the user states a fact, preference, decision, or project detail about themselves, or asks you to remember something. |
memory_feedback | write | Label a recall decision. Label a gated/verified recall decision right or wrong β only on the user's explicit confirmation; unlocks per-tenant self-learning. |
playbook_write | write | Store a proven strategy (trigger β what worked); deduped + credit-scored nightly. |
memory_write is a dedup'd bitemporal append β it never destroys or overwrites prior facts.
Full parameter-level reference: https://ultramemory.io/docs/tools/
Start in one click β connect UltraMemory with OAuth on Claude, ChatGPT, or Perplexity. No keys, no setup. The hosted server speaks OAuth 2.1 (PKCE) end-to-end, so the browser-based clients sign in without an API key; the terminal/CLI clients further down drive the same endpoint with an um_ key.
Endpoint: https://api.ultramemory.us/mcp (Streamable HTTP) Β· Auth: OAuth 2.1 (PKCE) for the browser connectors, or Authorization: Bearer um_<key> for CLI clients.
https://api.ultramemory.us/mcp β sign in when prompted. No terminal, no rule file.https://api.ultramemory.us/mcp β Auth = API key or OAuth. Read (recall/search) works on Plus/Pro developer mode; writes worked in our testing, but OpenAI's connector docs are in flux and conflict on write support there, so treat write as best-effort on Plus/Pro. Business/Enterprise/Edu workspaces get full read + write officially. Model note: the Instant model works with MCP; the Pro reasoning model currently disables MCP.UltraMemory β MCP server URL https://api.ultramemory.us/mcp β Advanced: OAuth (leave Client ID/Secret blank β dynamic registration) β Add β Connect (OAuth consent). Recall runs in Search mode; writes run in Computer mode β mention @UltraMemory to bind the connector. Verified end-to-end July 2026. Not available on Free; no marketplace submission yet β the connector is user-pasted.
Recommended β profile instructions: Settings β Personalization β Custom instructions, paste:
Perplexity caps Custom instructions at 1,500 characters β this text is 1,455 and fits; if you add your own lines, keep the total under 1,500 or the field silently truncates.
Terminal/CLI clients (Claude Code, Gemini CLI, Cursor, Codex, Windsurf, Cline, OpenClaw, VS Code): use the one-paste installs in Install options.
Claude Desktop (mcp-remote bridge, key instead of OAuth):
Hermes: see Hermes deep integration.
curl / REST:
Two ways to bring UltraMemory to a tool; you can start with the first and graduate to the second:
Want to stop burning tokens? The UltraMemory Plugin (one-line install) cut token use ~70% in our testing. Want that PLUS your project locked on persistent grounded truth β fewer iterations, faster delivery, and no tokens wasted on drift? Add the UltraMemory Agent Kit. Results may vary.
A Plugin (some platforms call it an extension) is a one-click bundle that packages several UltraMemory pieces into a single install: the memory connector (the nine tools), the recall-first rule/skill, and β on platforms that support them β the Turbo Token Saver hook and the checklist-bound-execution harness (skills + sub-agents). Instead of pasting a connector and a rule separately, you install one unit.
The token-saving hook that cut per-turn spend ~70% in our testing (measured 2026-07-05) and the harness sub-agents only run in agent runtimes that execute local hooks/sub-agents β Claude Code, Cowork, Hermes, and the Cline CLI. Cowork loads the connectors enabled on your claude.ai account (synced at session start) β add UltraMemory once in claude.ai, then toggle it on in Cowork's Customize sidebar. Browser chat clients (claude.ai, ChatGPT, Perplexity) run the connector's tools and rules but do not run local hooks or sub-agents, so on those surfaces a Plugin's value is the bundled connector + rule, not the hook/harness. Results may vary. This content is informational and not a guarantee of outcome.
| Platform | Native bundle concept | How UltraMemory rides it |
|---|---|---|
| Claude Code / claude.ai | Plugins (.claude-plugin) | The UltraMemory Agent Kit plugin: /plugin marketplace add LogicLabsAI/ultramemory-mcp β /plugin install ultramemory-kit@ultramemory. Bundles MCP + recall-first hook + harness skills/sub-agents. |
| Cursor | Plugins (Rules/Skills/Subagents/Commands/MCP/Hooks) β Cursor Marketplace + cursor.directory | Connect the hosted MCP server now (Cursor may force an OAuth login and ignore a static bearer); a full Cursor plugin bundle mirrors the agent-kit. |
| Gemini CLI | extensions (gemini extensions, gemini-extension.json) | gemini extensions install https://github.com/LogicLabsAI/ultramemory-mcp then gemini extensions config ultramemory (or export ULTRAMEMORY_API_KEY) β this repo ships gemini-extension.json + GEMINI.md. |
| OpenAI Codex | Plugins (.codex-plugin/plugin.json) | Remote MCP connector in config.toml; a Codex plugin bundle mirrors the agent-kit. |
| VS Code | Agent plugins (preview) + Extensions | Remote MCP server in mcp.json; VS Code also auto-detects the Claude plugin format. |
| Cline | Plugins (cline plugin install) β CLI/SDK/Kanban only today | Native plugin (plugins/cline/) installs via cline plugin install on Cline CLI/SDK/Kanban only β the VS Code + JetBrains extensions don't run plugins yet; those users use the connector/marketplace path. |
| OpenClaw | Plugins (native + Claude-compatible bundles) | Native MCP connector; can also install the Claude-format agent-kit bundle. |
| Hermes | Plugins + Skills (memory-provider kind) | The ultramemory-hermes memory-provider plugin β see Hermes deep integration. |
| Windsurf | No unified AI bundle β MCP servers + Rules + Workflows | Hosted remote MCP server + a .devin/rules/ recall-first rule (Windsurf's own "Plugins" are editor extensions, a different thing). |
| Perplexity | No unified plugin β Connectors (MCP) + Skills, installed separately; paid | Custom remote MCP connector + the recall-first skill: download skills/ultramemory-perplexity/SKILL.md, then Perplexity Computer β Skills β Create skill β Upload a skill (Pro/Max/Enterprise). |
The ultramemory-hermes package (this repo) is a full Hermes Agent memory provider β not just a
connector. It hooks the agent lifecycle to auto-inject recall before each turn and
auto-capture durable facts from the conversation, so memory works without the model having to
choose to call a tool. At session end it distills a whole-session rollup β both the user and
assistant sides are sent to the server, which curates one rich narrative card (blocker β approaches
β what worked β how verified); the per-turn sync_turn capture stays a raw turn record.
Install in three steps:
pip install ultramemory-mcp (also published as ultramemory-hermes)
If pip reports externally-managed-environment (PEP 668): pipx install ultramemory-mcp β or uv tool install ultramemory-mcp.ultramemory enable --key um_β¦ β writes the key to $HERMES_HOME/.env, plants the provider
shim at $HERMES_HOME/plugins/ultramemory/, and selects memory.provider: ultramemory in the
Hermes config.hermes memory status β verify the provider shows as installed.Hermes discovers memory providers by directory scan of $HERMES_HOME/plugins/ β it does not
consult Python entry points β so the shim planted by ultramemory enable is what makes the
pip-installed provider visible to Hermes. Setting environment variables alone CANNOT install the
provider: without ultramemory enable there is no shim on disk for the scan to find. To undo, run
ultramemory disable β it removes the shim and resets memory.provider to builtin.
On Teams, Business, and Enterprise accounts, memory is two-layer:
Recall blends both in one relevance-ranked query, so members automatically ground on company
knowledge plus their own context. In the Hermes provider, pick where auto-captured memory lands
with ULTRAMEMORY_SPACE:
ULTRAMEMORY_SPACE (choices private|shared, default private) sets the target space for
auto-writes (sync_turn, on_memory_write, on_session_end) and the default for the
memory_write tool. Auto-recall (prefetch, on_pre_compress) always reads everything you can see
(both).
The explicit tools also take an optional per-call space arg that overrides the default:
memory_write β space: private | shared.memory_recall / recall_gated β space: private | shared | both (default both).Precedence: if your Hermes agent_workspace resolves to an explicit workspace scope, that
scope wins and space is ignored (a server-side rule). space only takes effect for the default
(non-workspace) scope.
Within one account, the optional scope parameter partitions memory per project or workspace β
an explicit scope is written to and recalled from exclusively, so project A's memories never
bleed into project B:
scope='my-project' to UltraMemory tools."ULTRAMEMORY_SCOPE=my-project per project (see
hooks/README.md).Omit scope and everything shares the account default β one memory across all your tools, the
right default for personal use.
Want deterministic memory in Claude Code without Hermes? Two copy-paste, fail-open hooks:
UserPromptSubmit) β runs on every prompt you submit, recalls your top
matches, and injects them into context before the model answers.Stop) β runs when each turn finishes and sends the full turn (including tool
results) to UltraMemory, which distills the durable facts. Every Nth turn
(ULTRAMEMORY_SNAPSHOT_EVERY, default 5) it also nudges the model to author a wayback-grade
session snapshot via the bundled ultramemory-snapshot Skill
(Claude Code β₯ 2.1.163).Both are fail-open and copy-paste runnable. The copy-paste recall-hook install now lives in
Install options β Tier 2 above; full details (capture hook, global install,
per-project scopes) are in hooks/README.md.
The SDK clients in this repo (the Claude Code recall hook and the Hermes provider) opt into a
preview tier of recall that cuts per-turn token spend from thousands to hundreds, without
touching the hosted connectors β claude.ai, Claude Desktop, and ChatGPT behavior is unchanged
(the new mode / exclude_ids params are strictly opt-in; omitting them = full behavior).
mode: "preview": each non-policy fact renders as
a single line (- {fact_id} Β· {entity} Β· {key}: {first ~120 chars}β¦ (fetch for full)) under the
normal section headers, capped at ~2,000 chars. Full text stays one explicit fetch away.
[COMPANY POLICY] cards are exempt β they always render whole, in preview and full mode
alike (the anti-confabulation wedge is never truncated).exclude_ids, so
repeat turns don't re-spend budget on facts the model already holds; freed budget flows to fresh
facts.~/.ultramemory/cache.json (ejected by ultramemory enable; user-editable,
chmod 600, LRU-bounded at 500 entries / ~1 MB). It memoizes identical recall queries for 5
minutes (a repeat query makes zero HTTP calls) and tracks each session's seen fact_ids for
24 h. Delete the file to reset; corrupt files are silently rebuilt.Environment tunables:
| Env | Default | Effect |
|---|---|---|
ULTRAMEMORY_CACHE=off | on | kill switch β disables the memo + seen cache entirely |
ULTRAMEMORY_PREVIEW=off | on | Hermes prefetch reverts to full (non-preview) recall |
ULTRAMEMORY_HOOK_BUDGET | 2000 | Claude Code hook recall budget in characters |
ULTRAMEMORY_HOOK_POLICY_BUDGET | 12000 | hook injection cap on [COMPANY POLICY] turns (whole policies, never truncated) |
ULTRAMEMORY_MIN_CONFIDENCE | low | hook skips injection below this recall confidence |
CLAUDE.md rule for the agent's own mid-reasoning lookups, that trio is the real recall-first
guarantee.Apache-2.0 (see LICENSE). This is the open-source client surface. The UltraMemory
backend/engine β recall ranking, the metamemory gate, storage, metering, billing β is a separate,
proprietary hosted service at https://api.ultramemory.us.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/ultramemory)<a href="https://allmcps.com/mcp/ultramemory"><img src="https://allmcps.com/api/badge/ultramemory?style=directory" alt="UltraMemory on AllMCPs" /></a>