Local 2B and 0.8B decision models: calibrated yes/no, choice and score answers for your agent.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent — or use 1-click editor setup below.
One-click editor setup isn’t available for this listing yet — we don’t have a confirmed install command, and we’d rather show nothing than point your editor at the wrong package or host. Follow the project’s own setup instructions, linked above.
Small, calibrated decision models you run on your own machine, plus the tooling to put them to work in AI agents.
New in 0.3.0: Jev-Style-2B-Decision-v3. It scores 73.6 % on the 231 public items of JevBench v1.4.1 (self-run with the official harness on the GGUF F16 build, not an official board entry), 9.5 points above the 0.8B and the highest among the Qwen3.5-2B-family systems on the board. Its lead over decider-2b (71.0 %) is inside the 95 % confidence interval, 42 of the 82 board systems score higher, and hosted Jev is well ahead at 86.6 %. Try it in your browser, or run it locally with jev-style serve --release 2b. The default release is still the 0.8B, so existing setups get the same model as before.
Jev-Style is a family of small decision models built on Qwen3.5. The current releases are Jev-Style-2B-Decision-v3 (1.27 GB at 4-bit) and Jev-Style-0.8B-Decision-v3 (0.53 GB at 4-bit, the default); both, with every build and demo, are in the v3 collection. You give a model text or JSON and some typed questions, and it returns a calibrated probability for every option in one forward pass. The server's API follows the public systemone request shape, so clients written for Jev-compatible servers can call your laptop instead.
This repository is the part that makes the model useful day to day:
jev-style serve: a local API with a Playground and demos. It picks MLX on Apple silicon and PyTorch on CUDA or CPU; llama.cpp is optional.npx skills add. They serve the model, call it, evaluate it on your own labels, replace LLM calls that only return a label, add a guard to Claude Code, and add the MCP tools.PreToolUse hook where the local model checks every tool call before it runs and answers allow, ask or deny.decide, noul, choice and score for Claude Code, Cursor, Codex and any other MCP client.jev-style eval: measures accuracy and calibration on your own labelled data, and reports how many decisions you can automate at a 1, 5 or 10 % error budget.No GPU, no API key and no training needed.

noul), multiple choice (choice, up to 255 options) and ordered ratings (score, 2 to 10 levels). The model reads the text once and answers every question about it.Release (--release) | Parameters · smallest build | JevBench v1.4.1 public (231) | tweet_topic, zero-shot (1,693) | Context |
|---|---|---|---|---|
Jev-Style-2B-Decision-v3 (2b) | 1.9B · 1.27 GB (Q4_K_M) | 73.6 % | 82.2 % | 25,600 tokens |
Jev-Style-0.8B-Decision-v3 (0.8b, default) | 0.8B · 0.53 GB (Q4_K_M) | 64.1 % | 75.5 % | 25,600 tokens |
| Hosted Jev 1.13, for reference | – | 86.6 % | 79.3 % | – |
The 2B numbers are single pre-declared runs with its GGUF F16 build; its model card gives the protocols and confidence intervals, and the 2B Space runs it in the browser. On JevBench the 2B is the highest among the Qwen3.5-2B-family systems on the v1.4.1 board (decider-2b 71.0 %, open-jev-zefan-2b 64.5 %), but the lead over decider-2b is inside the 95 % confidence interval, 42 of the 82 board systems score higher, and hosted Jev is well ahead. On tweet_topic the 2B's accuracy is above Jev's published number, but its macro-F1 is below (0.678 vs 0.694). The 2B was trained on a reduced data pool (60M tokens) and has no separate limit for the question and its options; everything counts toward the 25,600 tokens.
The 0.8B against Laya, on sets neither was trained on:
| Model | Banking77 (77 intents, never trained) | MASSIVE intent, 37 held-out languages | tweet_topic, zero-shot | JevBench v1.4.1 public (231) |
|---|---|---|---|---|
| Jev-Style-0.8B-Decision-v3 | 68.2 % | 65.5 % | 75.5 % | 64.1 % |
| Best official Laya checkpoint (0.8B, 1,024 tokens by default) | 49.2 % | 36.1 % | 63.2 % | 58.4 % |
These numbers are from the 0.8B model card, which gives the full protocol and confidence intervals. The Laya rows are its official checkpoints re-run on the same rows, except tweet_topic and JevBench, which use published numbers. The hosted Jev API has higher accuracy than the 0.8B on every one of these sets where its accuracy is published. Treat both releases as small local options, not replacements for the hosted model.
| Build | 2B | 0.8B | Used by |
|---|---|---|---|
| safetensors: 2B, 0.8B | 3.76 GB | 1.50 GB | --backend torch (CUDA, Apple MPS, CPU) |
| MLX bf16 / 8-bit: 2B, 0.8B | 3.76 / 2.00 GB | 1.50 / 0.80 GB | --backend mlx (Apple silicon; auto picks it there) |
| GGUF F16 / Q8_0 / Q4_K_M: 2B, 0.8B | 3.78 / 2.01 / 1.27 GB | 1.52 / 0.81 / 0.53 GB | --backend gguf (llama.cpp through the release's scorer: jev-score-v2 for the 2B, jev-score for the 0.8B) |
Each build carries its own runtime file next to the weights. The server downloads a pinned revision and uses that file, so the answers here match what the model card documents. Stock llama.cpp, Ollama, LM Studio or mlx_lm.generate can load the weights but cannot produce the decision scores. The 2B MLX runtime needs mlx-lm 0.31.3 exactly, which jev-style[mlx] installs. For long documents on the 2B, use Q8_0 (the default) or F16 rather than Q4_K_M. Earlier 2B generations, for use in LM Studio or Ollama without this server: v1 GGUF (LM Studio, llama.cpp) and v2 GGUF (Ollama).
The 2B Space runs Jev-Style-2B-Decision-v3, and the 0.8B Space runs the 0.8B. There is nothing to install.
You'll need Python 3.10 or newer.
No reviews yet — be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/jev-style)<a href="https://allmcps.com/mcp/jev-style"><img src="https://allmcps.com/api/badge/jev-style?style=directory" alt="Jev Style on AllMCPs" /></a>