The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Arbitype listing page.
Typed decision tools for AI agents.
Classify · Score · Verify · Gate · Route · Review
MCP-native · Powered by TypeSafe Jev
Quick start · Why Arbitype? · Tools · Benchmarks · Documentation
Arbitype is an MCP-native typed decision layer for AI agents, powered by TypeSafe Jev. It turns probabilistic judgments into structured decision primitives that an agent or program can consume directly.
Arbitype is available on PyPI and the official MCP Registry.
[!NOTE] Arbitype is an independent open-source project. It is not an official TypeSafe AI product or an official integration for any particular agent host.
The fastest way to connect an MCP host is a local STDIO server launched by uvx:
For a pinned, reproducible launch:
Or install the package into the current environment:
The API key stays in the process environment. It is not an MCP argument and is never printed to standard output.
Inspect detected hosts before writing anything:
Then apply a reviewed plan:
The setup flow is plan → unified diff → confirmation → backup → apply. Use claude, cursor, or vscode in place of codex. Detailed host formats, secret forwarding, --yes, --remove, ownership, and non-interactive behavior are documented in docs/HOST_SETUP.md.
At minimum, set TYPESAFE_API_KEY. To select a model explicitly:
See the full configuration reference for endpoints, timeouts, retries, and request/response limits.
arbitype is the canonical distribution, Python package, and CLI. Existing TypeSafe MCP users should migrate to Arbitype; see docs/REGISTRY_MIGRATION.md.
For the safe migration from the historical distribution:
| Surface | Canonical | Legacy compatibility |
|---|---|---|
| PyPI / CLI | arbitype | typesafe-mcp, typesafe-codex-mcp |
| Python imports | arbitype | typesafe_mcp, typesafe_codex_mcp |
| MCP tools | route, review | codex_route, codex_review |
Arbitype is designed around contracts and bounded decisions rather than free-form prose:
| Principle | What it means |
|---|---|
| Contract-first | Requests and provider responses are validated before they reach the agent. |
| Safe by default | Credentials stay out of MCP tool arguments and generated host configuration. |
| Agent-oriented | Purpose-built classify, verify, gate, route, review, and score primitives. |
| Portable | Zero third-party runtime dependencies and standard MCP STDIO. |
| Measured | Public tool-selection and decision-stability evaluation assets. |
The decision flow is:
Start with one of the small, runnable fixtures:
| Scenario | Tool | What it demonstrates |
|---|---|---|
| Support routing | route | Choose one next action without executing it. |
| PR verification | verify | Check several claims independently. |
| Release review | review / gate | Separate holistic review from thresholded checks. |
| Agent next step | route | Keep the next workflow action bounded. |
The example outputs are illustrative fixtures. They are not live model results, accuracy claims, or authorization decisions.
Arbitype advertises nine read-only, idempotent MCP tools:
| Tool | Use when | Result |
|---|---|---|
| evaluate | You need a custom typed Jev question set. | Raw typed Jev response. |
| classify | You need one choice from unordered labels. | Choice and probability distribution. |
| score | You need one ordered rating or level. | Weighted score and distribution. |
| check | You need a probability for one bounded criterion. | Noul yes/no probability. |
| verify | You need several named claims checked independently. | Noul answer per claim. |
| gate | You need checks and thresholds transformed into a signal. | pass, review, or fail. |
| route | You need one suggested next action. | One action; no action execution. |
| review | You need a holistic quality or risk assessment. | Review decision and evidence. |
| health | You need local diagnostics or an explicit live check. | Configuration and optional provider health. |
| Need | Use | Do not substitute |
|---|---|---|
| One unordered label | classify | route, which selects an action |
| One ordered rating | score | classify, which has no order |
| One bounded proposition | check | verify, which handles multiple claims |
| Several named claims | verify | review, which assesses a whole object |
| Checks plus thresholds | gate | review, which is holistic |
| One next action | route | classify, which returns a category |
| Whole diff, plan, release, or report | review | verify, which answers claim by claim |
Tool descriptions also state USE WHEN and DO NOT USE WHEN boundaries so hosts can select a primitive without hidden prompt conventions.
Probabilities and confidence are model signals, not proof. gate and review are advisory decision transformations, not authorization systems, security boundaries, or approval engines.
The public product is Arbitype; TypeSafe Jev is the current provider. Provider configuration intentionally keeps the TYPESAFE_* names because the credential and endpoint belong to TypeSafe.
Recorded on the public 120-case Jev-mediated tool-selection evaluation:
| Metric | Result |
|---|---|
| Tool-selection accuracy | 96.67% |
| Invalid-tool rate | 0% |
| Schema-valid rate | 100% |
This is a Jev-mediated evaluation over Arbitype's advertised MCP tool catalog. It is not a Codex, Claude, or Cursor host benchmark and is not a general model-performance guarantee.
Details and reproducible assets:
Arbitype is a local MCP adapter and typed decision layer. It is not:
It does not execute actions suggested by route, edit files, run shell commands, or treat model probabilities as proof. Read SECURITY.md before using live credentials.
The test suite uses local fakes and does not need an API key:
The official MCP Python SDK interoperability smoke test is in scripts/official_sdk_smoke.py. A live Jev check is opt-in and paid:
Normal CI does not call the provider or run paid benchmarks.