The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Jev Style listing page.
Small, calibrated decision models you run on your own machine, plus the tooling to put them to work in AI agents.
New in 0.3.0: Jev-Style-2B-Decision-v3. It scores 73.6 % on the 231 public items of JevBench v1.4.1 (self-run with the official harness on the GGUF F16 build, not an official board entry), 9.5 points above the 0.8B and the highest among the Qwen3.5-2B-family systems on the board. Its lead over decider-2b (71.0 %) is inside the 95 % confidence interval, 42 of the 82 board systems score higher, and hosted Jev is well ahead at 86.6 %. Try it in your browser, or run it locally with jev-style serve --release 2b. The default release is still the 0.8B, so existing setups get the same model as before.
Jev-Style is a family of small decision models built on Qwen3.5. The current releases are Jev-Style-2B-Decision-v3 (1.27 GB at 4-bit) and Jev-Style-0.8B-Decision-v3 (0.53 GB at 4-bit, the default); both, with every build and demo, are in the v3 collection. You give a model text or JSON and some typed questions, and it returns a calibrated probability for every option in one forward pass. The server's API follows the public systemone request shape, so clients written for Jev-compatible servers can call your laptop instead.
This repository is the part that makes the model useful day to day:
jev-style serve: a local API with a Playground and demos. It picks MLX on Apple silicon and PyTorch on CUDA or CPU; llama.cpp is optional.npx skills add. They serve the model, call it, evaluate it on your own labels, replace LLM calls that only return a label, add a guard to Claude Code, and add the MCP tools.PreToolUse hook where the local model checks every tool call before it runs and answers allow, ask or deny.decide, noul, choice and score for Claude Code, Cursor, Codex and any other MCP client.jev-style eval: measures accuracy and calibration on your own labelled data, and reports how many decisions you can automate at a 1, 5 or 10 % error budget.No GPU, no API key and no training needed.

noul), multiple choice (choice, up to 255 options) and ordered ratings (score, 2 to 10 levels). The model reads the text once and answers every question about it.Release (--release) | Parameters · smallest build | JevBench v1.4.1 public (231) | tweet_topic, zero-shot (1,693) | Context |
|---|---|---|---|---|
Jev-Style-2B-Decision-v3 (2b) | 1.9B · 1.27 GB (Q4_K_M) | 73.6 % | 82.2 % | 25,600 tokens |
Jev-Style-0.8B-Decision-v3 (0.8b, default) | 0.8B · 0.53 GB (Q4_K_M) | 64.1 % | 75.5 % | 25,600 tokens |
| Hosted Jev 1.13, for reference | – | 86.6 % | 79.3 % | – |
The 2B numbers are single pre-declared runs with its GGUF F16 build; its model card gives the protocols and confidence intervals, and the 2B Space runs it in the browser. On JevBench the 2B is the highest among the Qwen3.5-2B-family systems on the v1.4.1 board (decider-2b 71.0 %, open-jev-zefan-2b 64.5 %), but the lead over decider-2b is inside the 95 % confidence interval, 42 of the 82 board systems score higher, and hosted Jev is well ahead. On tweet_topic the 2B's accuracy is above Jev's published number, but its macro-F1 is below (0.678 vs 0.694). The 2B was trained on a reduced data pool (60M tokens) and has no separate limit for the question and its options; everything counts toward the 25,600 tokens.
The 0.8B against Laya, on sets neither was trained on:
| Model | Banking77 (77 intents, never trained) | MASSIVE intent, 37 held-out languages | tweet_topic, zero-shot | JevBench v1.4.1 public (231) |
|---|---|---|---|---|
| Jev-Style-0.8B-Decision-v3 | 68.2 % | 65.5 % | 75.5 % | 64.1 % |
| Best official Laya checkpoint (0.8B, 1,024 tokens by default) | 49.2 % | 36.1 % | 63.2 % | 58.4 % |
These numbers are from the 0.8B model card, which gives the full protocol and confidence intervals. The Laya rows are its official checkpoints re-run on the same rows, except tweet_topic and JevBench, which use published numbers. The hosted Jev API has higher accuracy than the 0.8B on every one of these sets where its accuracy is published. Treat both releases as small local options, not replacements for the hosted model.
| Build | 2B | 0.8B | Used by |
|---|---|---|---|
| safetensors: 2B, 0.8B | 3.76 GB | 1.50 GB | --backend torch (CUDA, Apple MPS, CPU) |
| MLX bf16 / 8-bit: 2B, 0.8B | 3.76 / 2.00 GB | 1.50 / 0.80 GB | --backend mlx (Apple silicon; auto picks it there) |
| GGUF F16 / Q8_0 / Q4_K_M: 2B, 0.8B | 3.78 / 2.01 / 1.27 GB | 1.52 / 0.81 / 0.53 GB | --backend gguf (llama.cpp through the release's scorer: jev-score-v2 for the 2B, jev-score for the 0.8B) |
Each build carries its own runtime file next to the weights. The server downloads a pinned revision and uses that file, so the answers here match what the model card documents. Stock llama.cpp, Ollama, LM Studio or mlx_lm.generate can load the weights but cannot produce the decision scores. The 2B MLX runtime needs mlx-lm 0.31.3 exactly, which jev-style[mlx] installs. For long documents on the 2B, use Q8_0 (the default) or F16 rather than Q4_K_M. Earlier 2B generations, for use in LM Studio or Ollama without this server: v1 GGUF (LM Studio, llama.cpp) and v2 GGUF (Ollama).
The 2B Space runs Jev-Style-2B-Decision-v3, and the 0.8B Space runs the 0.8B. There is nothing to install.
You'll need Python 3.10 or newer.
--release works the same for decide, download and mcp --model; JEV_STYLE_RELEASE=2b sets it for every command and for the Python client.
To install the command line tool on its own, use uv: uv tool install "jev-style[all]" on Apple silicon, uv tool install "jev-style[torch]" elsewhere. The MCP server is part of every install since 0.3.0 (jev-style[mcp] still works).
Open http://127.0.0.1:8765 for the Playground and the demos: agent action approval, Snake, and Chinese and 51 languages. In another terminal, send a ticket:
This is the actual response from the 0.8B (the default release) with MLX on an Apple M1 Max, with probabilities rounded:
The ticket raises both a billing problem and a late delivery, and the probabilities show it: billing 0.64, shipping 0.29. That is why the model returns probabilities rather than a single label. Your code can act on the confident answers and hand the rest to a person.
For a script, one line is enough. The first call loads the model in-process (the 0.8B, or the release in JEV_STYLE_RELEASE), or uses the server in JEV_STYLE_URL if you set it:
For several questions about one text, and to choose the engine yourself:
from_pretrained takes the main repo (PyTorch), -MLX (precision="8bit" for the 2.00 GB 2B or 0.80 GB 0.8B weights) or -GGUF (quant="Q4_K_M", needs the release's scorer, jev-score-v2 for the 2B and jev-score for the 0.8B, see Backends). The client is not tied to this model: JevStyle(base_url=...) works with any server that implements POST /v1/systemone, and a client written for another systemone-compatible server can call http://127.0.0.1:8765 with any API key string.
The skills live in skills/. Each one tells a coding agent how to do one job from start to finish, and ends with a check that the job worked.
| Skill | What your agent does with it |
|---|---|
jev-style-serve | Installs the CLI, picks the backend for your machine, starts the server, checks it with a real request, and can optionally start it at login. |
jev-style | Writes code that calls the model: request format, reading the probabilities, turning them into actions with thresholds, limits. |
jev-style-eval | Builds a labelled JSONL from your data and runs jev-style eval. It reports accuracy, Brier, ECE, how much you can automate at a 1, 5 or 10 % error budget, and a refitted temperature. |
jev-style-adopt | Scans your code for LLM calls that only return a label, yes/no or a rating, and rewrites them as typed questions. The LLM stays as the fallback below a confidence threshold. It runs in shadow mode, measures agreement, then switches. |
jev-style-guard | Installs the Claude Code guard hook. It recommends a dry run first, then shows how to tune thresholds and measure them on labelled calls. |
jev-style-mcp | Registers the MCP tools with Claude Code, Codex, Cursor, Claude Desktop or Windsurf, and checks that a tool call works. |
In Claude Code you can also install everything as plugins:
Before Claude Code runs a Bash, Write, Edit, Read, WebFetch or MCP call, jev-style guard sends the call to the local model. It asks whether the call is destructive, exfiltrates data, touches secrets or goes outside the project, plus a 0 to 4 risk score. It turns the answers into allow / ask / deny. Regex hard rules in code can only make a verdict stricter. If the server is down, times out or rejects the request, the verdict is ask, never a silent allow.
On the 49 bundled tool calls, which were written and labelled by hand, the default config agrees with the labels on 77.6 %. No call labelled deny was allowed. 2 of the 16 calls labelled ask were allowed. The model alone, with no hard rules, agrees on 61.2 %. Use it as a second line of defence, not as a sandbox. The skill covers installing, dry runs and tuning.
jev-style mcp is a thin stdio server that forwards to the running jev-style serve, so one copy of the model serves every client. It is also listed in the official MCP Registry as io.github.lawrence3699/jev-style (runs as uvx jev-style mcp; set JEV_STYLE_URL if your server is not on http://127.0.0.1:8765). It provides decide (several questions about one input), noul, choice, score and model_info. Setup for Codex, Cursor and Claude Desktop is in the skill.
The best evidence is your own labelled examples. jev-style eval takes JSONL in the request format, with a label on each question:
The 30 example tickets were written by hand for this repository. They demonstrate the format; they are not a benchmark. Read automate @5%: 93.3% (p>=0.57) like this: if you act only when the top probability is at least 0.57, the model handles 93.3 % of tickets with at most 5 % errors among them, and everything else goes to a person or an LLM. The report also refits a temperature on half the rows and shows its effect on the other half. See the skill for building a proper evaluation set.
Repeat --server to run the same file through several engines: this package's local model, another build or quantisation, or any other server that implements POST /v1/systemone. The first one is the reference.
That run compares the MLX bf16 weights in-process with jev-style serve --precision 8bit on an Apple M1 Max. Engine specs are NAME=URL, NAME=local[:backend], NAME=hf:<repo id> or NAME=fake; --key NAME=ENV_VAR sends a bearer token from an environment variable and --model NAME=MODEL sets the request's model field. Every engine is scored on the same questions: if one engine rejects a row (too long, too many options), that row is dropped for all of them and counted under failed rows. --json PATH writes every answer.
POST /v1/systemone with {"state": text | object | array, "questions": {id: question}}:
type | criteria | Answer |
|---|---|---|
noul | optional {"true": "...", "false": "..."} | noul = P(true) |
choice | {option: description or null}, 1 to 255 options | choice, probabilities, confidence |
score | 2 to 10 levels, lowest first: "label" or {"label", "description"} | score = expected level index, legend, probabilities, confidence |
confidence = (k · p_max − 1) / (k − 1) for k options: 0 when the probabilities are uniform, 1 when one option takes all of them. The other routes are GET /v1/models, GET /healthz and the Playground at /. Errors look like {"error": {"code", "message", "question"?}} and use HTTP 422 (invalid_json, invalid_request, invalid_question, input_budget_exceeded), 401 (unauthorized), 404 or 500. Start the server with --api-key-env NAME to require Authorization: Bearer <key>. The full reference is skills/jev-style/reference.md.
| Machine | Command |
|---|---|
| Apple silicon | jev-style serve (MLX bf16; add --precision 8bit for 0.8 GB, or --release 2b --precision 8bit for the 2B at 2.0 GB) |
| NVIDIA GPU | jev-style serve --backend torch |
| CPU only | jev-style serve --backend torch --device cpu |
| llama.cpp | build the release's scorer once with sh build_jev_score.sh /path/to/llama.cpp from its GGUF repo (jev-score for the 0.8B, jev-score-v2 for the 2B), then jev-style serve --backend gguf --scorer /absolute/path/printed/by/the/script --quant Q4_K_M (add --release 2b for the 2B) |
| Offline | jev-style serve --model-dir /path/to/a/downloaded/model/repo |
| No model (UI or plumbing work) | jev-style serve --fake (deterministic, meaningless answers) |
Code: Apache-2.0 (LICENSE). The weights are Apache-2.0 fine-tunes of Qwen3.5-2B and Qwen3.5-0.8B; the NOTICE in each model repository lists the changes. The typed-question convention follows Laya.
Not affiliated with, endorsed by or connected to TypeSafe or Jev. "Jev-Style" describes the kind of model: a small typed-decision model in a similar style. No Jev weights, code or outputs are included. Not affiliated with Alibaba Cloud or the Qwen team or the Laya authors.