The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the B2IGE Verify listing page.
B2IGE Verify is deterministic verification infrastructure for AI-written software. It runs real programs under declared contracts, records evidence, and produces a bounded PASS/FAIL/INCONCLUSIVE/ERROR result without asking the coding agent to grade its own work. BlindTest makes the boundary concrete: hidden tests exercise the produced program without exposing the suite or oracle to the agent.
B2IGE Verify v0.3.0 is the current public release. The final release record documents its exact-main qualification, platform boundaries, and public asset set.
| Platform | Native archive |
|---|---|
| macOS Apple Silicon | b2ige-0.3.0-aarch64-apple-darwin.tar.gz |
| macOS Intel | b2ige-0.3.0-x86_64-apple-darwin.tar.gz |
| Linux x86_64 | b2ige-0.3.0-x86_64-unknown-linux-gnu.tar.gz |
Also download SHA256SUMS and verify the matching archive before extracting. See installation details for prerequisites and offline validation.
Choose the archive for the host, then verify, extract, and put bin on PATH:
The archive is ready for the adoption flow or the product-specific commands. Docker is required for BlindTest; missing prerequisites remain a readiness or non-PASS result.
BlindTest asks: does the produced program actually work? It runs visible checks against a correct and a buggy implementation, then executes sealed hidden cases in an attested Docker boundary. The correct implementation passes, the buggy implementation fails, and the sanitized Agent view is checked for private-value leakage.
From a source checkout or the source archive, with Docker available:
The demo performs real execution and records evidence; its output is not canned terminal text. See the BlindTest example for the bounded secrecy boundary and the threat model for its limits.
| Product | Question | What it checks |
|---|---|---|
| BlindTest | Does it actually work? | Sealed hidden tests against an agent-produced target inside the declared Docker boundary. |
| BehaviorSeal | Did it change? | Trusted reference/candidate behavior under equivalent deterministic experiments. |
| SideEffect Proof | Did it actually happen? | Committed local SQLite effects under retries and fault schedules, not inferred attempts. |
prepare is non-authoritative: it discovers inputs and writes a reviewable draft. Trust
approval remains an interactive human/trusted-controller decision. doctor reports readiness,
not verification. verify ID resolves the approved identity and produces the real,
evidence-backed result. Read the adoption guide for the full boundary.
These are committed measurements on the bounded B2IGE Verify Bench v1 corpus, not a claim of exhaustive correctness or a replacement for product evidence.
| Measurement (on the benchmark corpus) | Observed |
|---|---|
| Known bugs detected | 16 / 16 |
| False PASS | 0 |
| False FAIL | 0 |
| Verified reproduction | 16 / 16 |
| Hidden leakage observed | 0 |
| Agent private-value leakage | 0 |
| Mutation adequacy: known benchmark mutants | 2 / 2 |
All numbers above are on the benchmark corpus only. Bounded testing cannot establish complete correctness.
See the measured snapshot and benchmark methodology.
| Platform | Support boundary |
|---|---|
| macOS Apple Silicon | Native release |
| macOS Intel | Native release |
| Linux x86_64 | Native release; Docker and full benchmark CI |
| Windows x64 | Source/build plus bounded CLI smoke only; not VERIFIED_NATIVE |
| Linux arm64 | Deferred |
macOS binaries are unsigned and unnotarized. See installation and platform details for Docker, source-build, and archive boundaries.
Missing evidence cannot become PASS. Models do not assign verdicts. A result is bounded by the declared experiment, evidence, and platform scope.
Read the threat model, contracts, and evidence model before relying on a result.
Use the documentation index for getting started, product guides, trust and architecture, benchmarks, release history, and developer resources. The false PASS report accepts only synthetic or sanitized details; never upload hidden suites, private canaries, raw human evidence, or credentials.
Start with CONTRIBUTING.md, the architecture, and the security policy. Changes to verifier authority, evidence, contracts, or schemas require deliberate review; missing evidence must never be converted into success.
B2IGE Verify is licensed under Apache-2.0. No telemetry, hosted service, or paid API is required. Product names and trademarks are covered separately by TRADEMARKS.md.