Token-saving persistent memory for AI agents: recall durable facts instead of re-reading files.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
One-click editor setup isnβt available for this listing yet β we donβt have a confirmed install command, and weβd rather show nothing than point your editor at the wrong package or host. Follow the projectβs own setup instructions, linked above.
A framework for giving any AI agent persistent, browser-trainable SSM memory. It is provider-neutral β swappable LLM bridges (Anthropic / OpenAI / Fetch), an inference router, online distillation, and a persistent memory store β so any agent or app can depend on it, not just BuilderForce.
BuilderForce.ai is the flagship deployment, not the boundary: it trains a custom SSM in the browser, exports a model, and pushes it to an agent runtime where it runs as Evermind β the model itself, not just memory bolted onto someone else's LLM. Evermind is the whole brain: its own shared-expert generator (the cortex that does reasoning and language), a write-through knowledge memory (the hippocampus), and a trainable affective layer (the limbic system). A request can be served by Evermind directly β evermind/<ref> traffic is generated on-device, not forwarded to Claude/GPT β and external frontier models stay an optional routing choice rather than a hard dependency. The same framework is reusable by other agents off the shelf.
This monorepo consolidates two packages that previously lived in separate repos (mambacode.js and ssmjs) whose names described the technique rather than the role. They are renamed and unified here so that "what they do" is legible: they are Agent Memory.
| Package | Layer | Was | Responsibility |
|---|---|---|---|
@seanhogg/builderforce-memory-engine | Engine | @seanhogg/mambacode.js | WGSL/WebGPU Mamba SSM kernels, model blocks (Mamba1/2/3 + attention), autograd, trainer, BPE tokenizer, quantization. Zero runtime deps. |
@seanhogg/builderforce-memory | Runtime | @seanhogg/ssmjs | SSM execution, Transformer orchestration (Anthropic/OpenAI/Fetch bridges), online distillation, inference router, sessions, the persistent MemoryStore, and Write-Through Cognition (EvermindCognition). Depends on memory-engine. |
The two-package split is deliberate: the engine is zero-dep and WebGPU-pure and can be consumed standalone; the runtime pulls in LLM-vendor bridges. Flattening them would force engine-only consumers to drag in vendor code and vice versa. They release in lockstep from one pipeline, which kills the publish-drift bug class that the separate-repo setup suffered (bumping one version without publishing + regenerating the consumer lockfile).
Agent Memory ships three layers that reduce LLM spend, in increasing power. All are portable β the same code runs in the browser (WebGPU SSM) and in Node (the agent's @webgpu/node SSM) because the embedder and storage are injected, never hard-wired.
| Layer | What it does | Saves |
|---|---|---|
Prompt caching (AnthropicBridge cacheSystem) | Marks the stable system prompt as an Anthropic cache_control block | ~90% on the cached input prefix (cost, not count) |
Exact-match cache (CachingBridge + ResponseCache) | Reuses byte-identical completions | Eliminates duplicate calls (retries, identical fan-out) |
Semantic cache (SemanticCache + SemanticCachingBridge) | Reuses a prior answer when the new prompt is within a cosine threshold of one already answered β catches paraphrases | Avoids frontier calls entirely on semantically-repeated prompts |
The semantic cache is the real lever. It is read-through with two tiers, mirroring an L1-Map / L2-KV cache:
FetchSemanticCacheBackend β the BuilderForce.ai gateway), so a paraphrase answered by the web app is reusable by an agent and vice-versa.Memory-backed fact injection is semantic too: SSMAgent defaults to factSelection: 'semantic', injecting only the top-maxFacts embedding-relevant facts each turn (paraphrase-robust, smaller prompts) instead of an exact key-substring match.
Caching keeps answers fresh; cognition keeps knowledge fresh. A frozen model β or an append-only memory β drifts: stale and current facts pile up under different keys until something reconciles them by hand. EvermindCognition closes that gap. It is the model-knowledge analogue of a write-through cache with a conflict resolver, so a belief is replaced on write, never appended into a reconciliation backlog. This is what lets Evermind's knowledge stay current without a manual reconcile step.
Every candidate fact flows through one pipeline:
Canonicalize (stable subject key) β recall incumbent β evaluate evidence β reconcile (augment | confirm | supersede | reject) β write-through
EvidenceGatherer is injected, so the evidence source is surface-specific (IDE file tools, a DB probe, an HTTP check); workspacePresenceGatherer ships for the common "is it still on disk?" case.recall() is served from a version-token cache that invalidates the instant knowledge changes β the same invalidate-on-write rule as the semantic cache, applied to beliefs.EvermindCognition is store- and surface-agnostic β the CognitionFactStore interface is satisfied structurally by MemoryStore (no adapter) β so the same loop runs in the IDE, on-prem, cloud, and the browser.
Caching cuts the bill and cognition keeps knowledge current, but neither answers the questions an enterprise buyer actually asks before signing: can it read our data, will it leak across tenants, what did that answer cost, and how do we know it got better? Those four questions are what this layer exists to answer. Each is a port with adapters behind it, so a customer's existing stack is an adapter choice rather than a rewrite.
| Layer | Entry point | Subpath export | Answers |
|---|---|---|---|
| Ingestion | IngestionPipeline | @seanhogg/builderforce-memory/ingest | Can it read our data β structured and unstructured β and stay current? |
| Vector store | VectorStore port + adapters | β¦/vectorstore | Can it run against the database we already have, with our tenancy rules? |
| Retrieval | EnterpriseRetriever | β¦/rag | Are the answers grounded, scoped, and citable? |
| Orchestration | AgentGraph + patterns | β¦/orchestration | Can agents coordinate, pause for a human, and resume after a crash? |
| Telemetry | Tracer + MetricsRegistry | β¦/telemetry | What did it cost, how fast was it, and where did it go wrong? |
| Evaluation | EvalHarness | β¦/eval | How do we know it is good enough to launch β and still is? |
A support-ticket export (rows, typed columns) and a policy document (prose) reach the index through the same pipeline; the difference is which parser ran, not which system was built. Structure that survives parsing becomes filterable metadata, which is what makes "summarise open P1 tickets about billing" answerable β status and priority are database predicates rather than something the embedding has to imply.
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/builderforce-memory)<a href="https://allmcps.com/mcp/builderforce-memory"><img src="https://allmcps.com/api/badge/builderforce-memory?style=directory" alt="Builderforce Memory on AllMCPs" /></a>