The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Factum listing page.
Status: v0.1.3 — Early stage, seeking early collaborators. Core write/query/retract pipeline works. Not production-ready. Architectural decisions are still open to change.

Interactive docs — animated syntax parsing, 7-tuple explorer, query pipeline, token efficiency chart, and MCP architecture diagram.
Every fact an agent writes carries mandatory provenance. When a source is retracted, everything derived from it is invalidated automatically — cascade retraction. When facts conflict, Factum returns Ambiguous instead of guessing.
A structured knowledge language (S-expression based), Rust implementation, stdio MCP server — works with Claude Code, Cursor, and any MCP client.
Status: v0.1.3, early stage. Core write/query/retract pipeline works; no semantic search or memory consolidation yet — Factum handles verified structured facts, not conversation context. Best suited for compliance-sensitive agents, multi-agent shared knowledge bases, and anywhere "why did the agent believe X" needs an answer. See the multi-agent usage guide for shared knowledge base patterns.
Unique to Factum: grammar-enforced provenance · cascade retraction · conflict refusal.
Not (yet) in Factum: embedding-based semantic retrieval · memory consolidation · HTTP transport.
Complementary: Mem0/Letta store and retrieve context; Factum stores auditable structured facts. They can run side by side via MCP.
| Agent memory pain point | Factum mechanism |
|---|---|
| Can't tell "user said" from "LLM inferred" | 5-level provenance + grammar-enforced model name on Extracted nodes |
| Stale memory used as current fact | Soft delete + cascade retraction via reverse dependency graph |
| Conflicting memories silently pick one | Ambiguous — refuses to answer rather than guess |
| Memory pollution (prompt injection) | Provenance chain makes contamination traceable and retractable |
| Enterprise can't let agents store sensitive data | Index-level permission filtering — no post-query leakage |
Academic context: The STALE benchmark (2025) shows that even the best LLM agents achieve only 55.2% accuracy at detecting when their own memories are outdated — confirming that memory staleness is an unsolved problem in agent systems.
:note for human-readable context"auto" as the node ID in factum_insert or factum_assert to auto-generate content-based IDs — no manual ID managementExtracted nodes MUST carry model + version — the parser rejects them if missing (not just a documentation convention)deps_rev reverse dependency graph); configurable node limit prevents cascade explosion in large knowledge basesfactum_assert returns corroborated (not an error) — independent agreement is the most valuable multi-agent signalAmbiguous when it cannot uniquely resolveDec(i128, u8) — zero floating-point error| Module | Status | Notes |
|---|---|---|
| factum-core (types, lexer, parser, serialize) | ✅ Implemented | 100% syntactic round-trip |
| factum-rt (store, query, arbitration, permissions, verifiers) | ✅ Implemented | In-memory store (default) + RocksDB persistence (--features rocksdb); 5 column families, WriteBatch atomic writes |
| factum-mcp (JSON-RPC bridge) | ✅ Protocol + handler | stdio transport verified end-to-end; HTTP transport not yet implemented (remote deployment requires custom wrapper) |
| factum-bench (benchmarks) | ✅ Implemented | Syntax round-trip + token efficiency + query perf |
| factum-l (latent space projection) | ❌ Not started | Research item — see ROADMAP.md |
| Wikidata/Mathlib corpus converters | ❌ Not started | M2 milestone |
| Lean/Z3 verifiers | ❌ Not started | Only Schema + DecimalRange verifiers implemented |
| Embedding-based semantic search | ❌ Not started | Not on roadmap — consider using Mem0 alongside Factum |
| Memory consolidation/summarization | ❌ Not started | Not on roadmap — consider using Letta alongside Factum |
| Inspector (visual debugger) | ❌ Not started |
See Getting Started with Factum MCP for the full guide.
When n001 is retracted, n006 is automatically invalidated — the agent knows it can no longer trust the subsidiary relationship.
| Layer | What It Means | Status |
|---|---|---|
| Agent Read | Agent receives Factum-F as context via MCP — lower token overhead than verbose JSON | ✅ Token efficiency measured (real o200k_base: canonical −62%, compact −53% vs JSON) |
| Agent Write | Agent generates Factum-F nodes — parse uniqueness guarantees one valid interpretation, error classes enable self-correction | ✅ See authoring guide |
| Latent Reasoning | factum-l: encode Factum-F into continuous thought vector, reason in latent space, decode back for audit | 🔬 Research item — not started, not blocking layers 1-2 |
A structured knowledge language underpins the memory layer — S-expression based for parse uniqueness, with Dec(i128, u8) for lossless numerics and mandatory model references for LLM self-auditing. Full design rationale in docs/design-rationale.md.
TL;DR: Factum canonical form saves ~62% tokens vs verbose JSON with the same metadata. All numbers measured with real o200k_base (GPT-4o) tokenizer via tiktoken-rs.
| Format | Real tokens (5 nodes) | vs verbose JSON | What it includes |
|---|---|---|---|
| Factum canonical | 238 | −62% | Full 7-tuple: provenance + confidence + validity + permissions |
| Factum compact (JSON) | 290 | −53% | Same 7-tuple, JSON with numeric tags |
| Markdown | 181 | −71% | Assertion text only — no provenance, no confidence, no permissions |
| JSON (pretty) | 623 | baseline | Same 7-tuple metadata in verbose JSON encoding |
Run
cargo test -p factum-bench test_token_efficiency_real_tokenizer -- --nocaptureto reproduce.
Key finding: The canonical S-expression form (238 tokens) is more token-efficient than the compact JSON form (290 tokens) — BPE tokenizers split JSON delimiters but merge S-expression parentheses. The form designed for correctness is also the most token-efficient.
Syntactic round-trip ✅ — parse(serialize(parse(x))) == parse(x). 100% and verified by the test suite + fuzzing. This is the trust foundation.
Semantic round-trip ❌ Not yet measured — requires factum-l (latent space projection, not implemented).
morphemes.toml via build.rs)morphemes.toml codegen is a future M2 item for external contributions.| Format | How Factum Differs |
|---|---|
| RDF / JSON-LD | RDF triples carry no per-node provenance, confidence, or permissions. Factum makes these first-class and non-optional. |
| Markdown | Markdown has zero metadata. Factum trades human readability for machine verifiability. |
| JSON | JSON has no schema, no provenance, no temporal validity. Factum compact form uses JSON as transport but adds structure and audit chain. |
| Mem0 / Zep / Letta | These store and retrieve agent memory. Factum adds grammar-enforced provenance, cascade retraction, and conflict refusal. Complementary, not competitive. |
Run cargo test to see the current count. Tests cover: types, lexer, parser, serialization, morphemes, depth-limit/DoS protection, conformance vectors, store, query, arbitration, permission, verifier, subscription, MCP protocol/tools/handler, and benchmarks.
CI badge at the top of this README reflects the latest build status.
See SECURITY.md for vulnerability disclosure.
MIT
"Factum" and the Factum-F language specification are project names. This MIT license covers code only; the specification may be governed by a separate process in the future.
Contributions require DCO sign-off (git commit -s). See CONTRIBUTING.md.
mcp-name: io.github.factum-project/factum