The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Vaara listing page.
Accountable Autonomy.
A verifiable receipt for every autonomous action, checkable by anyone.
Your AI agent transferred the funds, wrote the file, called the tool. Later, someone who does not trust you asks you to prove exactly what it did and why: a regulator, an auditor, a customer after an incident. Your own logs will not settle it, because you could have edited them.
That is the whole thing. Every call to a governed function is risk-scored and decided against your policy before the body runs. An allowed call runs, and the decision, the call, and the outcome land in a hash-chained, tamper-evident record anyone can verify offline. Sign it at export (vaara trail export) for third-party proof. Records persist to ~/.vaara/trail/audit.db by default, so evidence survives restarts. Python 3.10+, zero runtime dependencies.
Both deny and escalate raise vaara.Blocked, since an escalation means a human has not answered yet. Run the example above on a fresh install and it will raise: with no outcome history the scorer's confidence interval is wide, and a tx.transfer escalates on the interval's upper bound even though its point estimate sits under the allow threshold. That is the intended direction to fail, and it settles. Feeding real outcomes back through report_outcome narrows the interval, and the same call starts allowing after a few dozen clean results. To watch decisions without acting on them while that happens, start with @vaara.govern(shadow=True).
vaara.io/verify.html is the Vaara Resin. One HTML file, no build step and no dependencies. Paste in a receipt and it recomputes the DSSE pre-authentication encoding, takes its digest, and checks the Ed25519 signature with WebCrypto. The receipt never leaves the tab, nothing uploads, and the page works with the network off, so verification is not a service and Vaara is not a party to it. Save the file and it keeps working.
It also states what a passing check does not establish: that the key belongs to the party you expect, that the signed statement is true, that decided_at means anything without an external time authority, or that one receipt is a whole history.
The explorer on the same page reads the public transparency log straight from your browser. Look a trail head up by digest, or paste a public key to see everything published under it. No account and no sign-in, because the key is the identity. Publishing to that log is opt-in and off by default (vaara trail publish-head), so an absence there means nothing was published rather than nothing happened.
vaara.io/conformance.html is the results page. It carries every suite and its verdict, and every party other than the maintainer who ran the checkers and reported what they found in public. Rows are chained, each holding the digest of the row before it, so removing or reordering one breaks every digest after it and the break is visible to anyone. The maintainer cannot take a row down either. A run that disagrees with ours is a row too, with the reason stated, and there is no blacklist.
The aggregate runner grades every suite at once, and grades another implementation's vectors the same way:
It prints a prefilled link at the end of every run, so asking for a row takes one click. The named, versioned rule set, what a pass does and does not establish, and the full suite list are in docs/conformance-profile.md.
The decorator drives the same engine you can call directly when you want the decision object in hand.
Every call gets a risk score and an allow / block / escalate decision against your policy, then the call, the decision, and the real outcome are written to the audit trail. report_outcome closes the loop: the scorer reweights based on which signals actually predicted the outcome. Releases ship SLSA Build Level 3 provenance, verifiable with slsa-verifier verify-artifact. Optional ML classifier: pip install 'vaara[ml]'.
Writing a trail is the easy half. The half that matters is letting someone who does not trust you check it, with no key, no access, and none of your code. Every Vaara record is content-addressed and fail-closed on authenticity, and ships with public conformance vectors plus a standalone checker that imports no Vaara code, so an independent party reproduces every verdict offline.
ok only when a signature is actually established, not merely present in a log. The same property drives the standards work behind the Vaara Receipt Internet-Draft: evidence that holds up for someone who runs none of your software. The full verifier set, the trust model for each verb, and where trust comes from in each case are in docs/verifying-evidence.md.
To check that claim yourself, without installing Vaara, run the standalone checker against the published vectors. Its only dependencies are cryptography and rfc8785:
It re-derives every verdict from the receipt bytes and the public key alone. The output shows the property the trail is built for: a receipt dropped from inside a declared boundary is a provable gap from the held set, with no issuer access and no external witness.
For the whole loop in one runnable file, produce a signed record, verify it yourself, then watch a single forged byte get caught, see examples/prove-it-yourself/. The logs-versus-evidence argument behind it is in docs/logs-vs-evidence.md.
vaara compliance report --format json against a real trail produces an article-level evidence record an auditor reads directly. Articles with no recorded events return evidence_insufficient, not a rubber stamp.
Each verdict carries the threshold-versus-observed snapshot, the rationale, and the underlying records, so a reviewer traces status back to a concrete event. The same data renders as a Notified-Body PDF, a static HTML dashboard, or a Sigstore-signed handoff envelope. See docs/COMPLIANCE.md.
vaara verify-contiguity). Off by default.Native adapters route the major Python agent frameworks through the same pipeline, each via the framework's own hook, emitting identical audit events:
| Framework | Entry point |
|---|---|
| LangChain | VaaraCallbackHandler, vaara_wrap_tool |
| CrewAI | VaaraCrewGovernance |
| OpenAI Agents SDK | VaaraToolGuardrail, vaara_wrap_function |
| MCP server | vaara.integrations.mcp_server |
To put Vaara in front of an MCP server, run it as a proxy. Every tools/call routes through the pipeline before reaching the upstream; allowed calls forward transparently, blocked calls return an MCP error.
Start with --shadow: every call is classified, scored, and recorded, nothing is blocked. After a few days, vaara trail shadow-report --db ./mcp_audit.db shows what enforcement would have done; then drop the flag and enforce, starting from a ready-made perimeter for common MCP servers in examples/policies/mcp-starters/. Point your MCP client (Claude Code, Cursor, any host) at the proxy instead of the upstream. There is also an HTTP API (pip install 'vaara[server]', vaara serve) and a first-party TypeScript client on npm (@vaara/client) for non-Python agents. Framework details, the cloud and OSS guardrail adapters (Bedrock, Azure, GCP, NeMo, Guardrails AI, LLM Guard, Rebuff), and the multi-tenant proxy are in docs/adapters.md.
A policy is code, so it belongs in the pull request that changes it. The action validates the policy, runs its cases, and fails the build on a policy that does not parse, a failing case, or a trail whose chain or signature does not hold.
Point trail at a signed zip to verify one a job produced. Without a pubkey that checks the trail is internally intact; pass a key you obtained separately to bind the signer too, and the run says which of the two it did. Inputs, outputs and pinning are in docs/github-action.md.
This checks artifacts. Gating the agent is the runtime's job, at the moment of the tool call.
Each risk score blends five expert signals and keeps adapting as outcomes come back, and it carries a confidence interval with a coverage guarantee that holds regardless of the input distribution. On a held-out adversarial corpus the classifier reaches 84.7% recall (95% Wilson [82.4, 86.7]) at a 4.1% false-positive rate, and 1.2% FPR on benign calls under live injection pressure. The hot-path rule scorer adds 140 µs mean per call on commodity CPU; the ML classifier is opt-in (vaara[ml]) and off that path. make bench reproduces the classifier figures below against the bundles that ship in the repo; it needs pip install 'vaara[ml]' and downloads the embedding model on first run.
Method and per-cell breakdown: docs/architecture.md and bench/.
draft-sirkkavaara-vaara-receipt. A second independent implementation has reproduced the SEP-2828 conformance vectors from a clean checkout with no shared code. Published corpora with independent checkers cover the fallback binding path (tests/vectors/fallback_projection_v0/) and CrewAI governance decisions (tests/vectors/governance_decision_v0/).Details and the offline checkers for each: docs/standards.md.
The public surface is fixed: the signed envelope (vaara.receipt/v1), capability constraints, the credential grant and gateway, and the @vaara.govern entry point. No new primitives are planned. New behavior ships as profiles that pin to vaara.receipt/v1, not as new core types, and no new format bindings will be added (the last was v1.13.0). From here the work is hardening and subtraction within this surface, so anyone building on it has a stable target.
| Path | Contents |
|---|---|
| docs/verifying-evidence.md | Every verifier and its trust model |
| docs/logs-vs-evidence.md | Logs vs evidence: proving what an agent did, and what the AI Act actually requires |
| docs/prove-what-an-ai-agent-did.md | The four properties a provable record of agent actions needs |
| docs/eu-ai-act-article-12.md | Article 12 record-keeping: what it requires, what it does not, what to demand from tooling |
| docs/tamper-evident-audit-trail.md | How the trail works, its honest limits, and what it costs |
| docs/vaara-vs-observability-vs-grc.md | Vaara vs Datadog/Splunk vs Vanta/Drata: three different questions |
| docs/dogfood/ | Our marketing runs under this gate; the signed trail and key to verify it |
| docs/architecture.md | Scoring, conformal coverage, time anchor, formal properties |
| SPEC.md | The canonical vaara.receipt/v1 receipt format spec |
| docs/standards.md | SEP-2828, SEP-2787, OVERT, the sovereign inference harness |
| docs/adapters.md | Framework and cloud/OSS guardrail adapters, multi-tenant proxy |
| docs/COMPLIANCE.md | EU AI Act and DORA article mapping, eval numbers |
| docs/multi-replica-deployment.md | Scaling past one proxy process: per-replica chains, rotation, archive index |
| docs/kubernetes-rancher.md | Running the proxy on Kubernetes with Rancher: chart, storage, enforcement, network isolation |
| docs/supported-platforms.md | Python, container and Kubernetes versions Vaara supports, and which have been verified |
| CHANGELOG.md | Version-by-version evolution |
| docs/PRIOR_ART.md | When each concept first shipped, plus adjacent work |
Vaara helps deployers assemble evidence for their own conformity work. It does not certify compliance or constitute legal advice. Deployers own their obligations under the EU AI Act and other applicable law.
Commercial license and paid pilots available: see vaara.io or contact hello@vaara.io. Licensing terms are in LICENSING.md, and the commercial licence is described in COMMERCIAL.md.
If you build on Vaara or its receipt format, cite the repository (see CITATION.cff) and the specification it implements:
Henri Sirkkavaara. The Vaara Receipt: A Recomputable Receipt Format for Decisions About Agent Actions. IETF Internet-Draft draft-sirkkavaara-vaara-receipt.
Every tagged release is archived by Zenodo and minted a DOI. Cite 10.5281/zenodo.22027975 for the software as a whole, which always resolves to the newest version, or the version DOI printed on a specific release for the exact bytes you ran.
Copyright © 2026 Henri Sirkkavaara. Licensed under AGPL-3.0-or-later. See LICENSE.