Knowledge language & MCP server: verifiable facts with provenance, confidence, audit-safe retraction
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
One-click editor setup isnβt available for this listing yet β we donβt have a confirmed install command, and weβd rather show nothing than point your editor at the wrong package or host. Follow the projectβs own setup instructions, linked above.
Status: v0.1.3 β Early stage, seeking early collaborators. Core write/query/retract pipeline works. Not production-ready. Architectural decisions are still open to change.

Interactive docs β animated syntax parsing, 7-tuple explorer, query pipeline, token efficiency chart, and MCP architecture diagram.
Every fact an agent writes carries mandatory provenance. When a source is retracted, everything derived from it is invalidated automatically β cascade retraction. When facts conflict, Factum returns Ambiguous instead of guessing.
A structured knowledge language (S-expression based), Rust implementation, stdio MCP server β works with Claude Code, Cursor, and any MCP client.
Status: v0.1.3, early stage. Core write/query/retract pipeline works; no semantic search or memory consolidation yet β Factum handles verified structured facts, not conversation context. Best suited for compliance-sensitive agents, multi-agent shared knowledge bases, and anywhere "why did the agent believe X" needs an answer. See the multi-agent usage guide for shared knowledge base patterns.
Unique to Factum: grammar-enforced provenance Β· cascade retraction Β· conflict refusal.
Not (yet) in Factum: embedding-based semantic retrieval Β· memory consolidation Β· HTTP transport.
Complementary: Mem0/Letta store and retrieve context; Factum stores auditable structured facts. They can run side by side via MCP.
| Agent memory pain point | Factum mechanism |
|---|---|
| Can't tell "user said" from "LLM inferred" | 5-level provenance + grammar-enforced model name on Extracted nodes |
| Stale memory used as current fact | Soft delete + cascade retraction via reverse dependency graph |
| Conflicting memories silently pick one | Ambiguous β refuses to answer rather than guess |
| Memory pollution (prompt injection) | Provenance chain makes contamination traceable and retractable |
| Enterprise can't let agents store sensitive data | Index-level permission filtering β no post-query leakage |
Academic context: The STALE benchmark (2025) shows that even the best LLM agents achieve only 55.2% accuracy at detecting when their own memories are outdated β confirming that memory staleness is an unsolved problem in agent systems.
:note for human-readable context"auto" as the node ID in factum_insert or factum_assert to auto-generate content-based IDs β no manual ID managementExtracted nodes MUST carry model + version β the parser rejects them if missing (not just a documentation convention)deps_rev reverse dependency graph); configurable node limit prevents cascade explosion in large knowledge basesfactum_assert returns corroborated (not an error) β independent agreement is the most valuable multi-agent signalAmbiguous when it cannot uniquely resolveDec(i128, u8) β zero floating-point error| Module | Status | Notes |
|---|---|---|
| factum-core (types, lexer, parser, serialize) | β Implemented | 100% syntactic round-trip |
| factum-rt (store, query, arbitration, permissions, verifiers) | β Implemented | In-memory store (default) + RocksDB persistence (--features rocksdb); 5 column families, WriteBatch atomic writes |
| factum-mcp (JSON-RPC bridge) | β Protocol + handler | stdio transport verified end-to-end; HTTP transport not yet implemented (remote deployment requires custom wrapper) |
| factum-bench (benchmarks) | β Implemented | Syntax round-trip + token efficiency + query perf |
| factum-l (latent space projection) | β Not started | Research item β see ROADMAP.md |
| Wikidata/Mathlib corpus converters | β Not started | M2 milestone |
| Lean/Z3 verifiers | β Not started | Only Schema + DecimalRange verifiers implemented |
| Embedding-based semantic search | β Not started | Not on roadmap β consider using Mem0 alongside Factum |
| Memory consolidation/summarization | β Not started | Not on roadmap β consider using Letta alongside Factum |
| Inspector (visual debugger) | β Not started |
See Getting Started with Factum MCP for the full guide.
When n001 is retracted, n006 is automatically invalidated β the agent knows it can no longer trust the subsidiary relationship.
| Layer | What It Means | Status |
|---|---|---|
| Agent Read | Agent receives Factum-F as context via MCP β lower token overhead than verbose JSON | β Token efficiency measured (real o200k_base: canonical β62%, compact β53% vs JSON) |
| Agent Write | Agent generates Factum-F nodes β parse uniqueness guarantees one valid interpretation, error classes enable self-correction | β See authoring guide |
| Latent Reasoning | factum-l: encode Factum-F into continuous thought vector, reason in latent space, decode back for audit | π¬ Research item β not started, not blocking layers 1-2 |
A structured knowledge language underpins the memory layer β S-expression based for parse uniqueness, with Dec(i128, u8) for lossless numerics and mandatory model references for LLM self-auditing. Full design rationale in docs/design-rationale.md.
TL;DR: Factum canonical form saves ~62% tokens vs verbose JSON with the same metadata. All numbers measured with real o200k_base (GPT-4o) tokenizer via tiktoken-rs.
| Format | Real tokens (5 nodes) | vs verbose JSON | What it includes |
|---|---|---|---|
| Factum canonical | 238 | β62% | Full 7-tuple: provenance + confidence + validity + permissions |
| Factum compact (JSON) | 290 | β53% | Same 7-tuple, JSON with numeric tags |
| Markdown | 181 | β71% | Assertion text only β no provenance, no confidence, no permissions |
| JSON (pretty) | 623 | baseline | Same 7-tuple metadata in verbose JSON encoding |
Run
cargo test -p factum-bench test_token_efficiency_real_tokenizer -- --nocaptureto reproduce.
Key finding: The canonical S-expression form (238 tokens) is more token-efficient than the compact JSON form (290 tokens) β BPE tokenizers split JSON delimiters but merge S-expression parentheses. The form designed for correctness is also the most token-efficient.
Syntactic round-trip β
β parse(serialize(parse(x))) == parse(x). 100% and verified by the test suite + fuzzing. This is the trust foundation.
Semantic round-trip β Not yet measured β requires factum-l (latent space projection, not implemented).
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/factum)<a href="https://allmcps.com/mcp/factum"><img src="https://allmcps.com/api/badge/factum?style=directory" alt="Factum on AllMCPs" /></a>