Jev judgments for agents: injection scanning, shell risk gating, ranking
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent — or use 1-click editor setup below.
One-click editor setup isn’t available for this listing yet — we don’t have a confirmed install command, and we’d rather show nothing than point your editor at the wrong package or host. Follow the project’s own setup instructions, linked above.
用 TypeSafe Jev(System One 决策模型)做 agent harness 工程实验的技术仓库:评测框架、基准报告、可运行的集成工具(MCP server、skill router)。
用 uv 一条命令,无需克隆:
配置到任意 MCP 客户端(opencode / Claude Desktop 等):
提供的工具:
| 工具 | 作用 |
|---|---|
scan_injection | 扫描工具/网页/邮件输出里的注入指令,返回 block / review / pass |
bash_risk | shell 命令四维风险打分(破坏性 / 触密 / 外发 / 不可逆),返回 deny / review / allow |
rank_candidates | 候选片段按相关性打分排序(RAG 精排) |
MCP Registry 归属标记(勿改格式):
mcp-name: io.github.Aitejiu/jev
用 Jev 把任务路由到已安装的 skill,并且只加载选中那一个的完整指令,避免把整个 skill 目录塞进主模型上下文。技能页:https://www.skills.sh/aitejiu/jev-harness-lab/jev-skill-router
两个集成都需要:
评测结果缓存在 eval/results/*.jsonl(已随仓库提交,可 --report-only 直接出报告);原始数据集在 eval/data/,需按 eval/README.md 的说明下载(已 gitignore)。
| 工具 | 作用 | 输入 → 输出 |
|---|---|---|
scan_injection | 扫描工具输出中的注入指令 | tool_output → action(block/review/pass) + 概率 |
bash_risk | shell 命令四维风险打分 | command → action(deny/review/allow) + 破坏性/触密/外发/不可逆分数 |
rank_candidates | 候选片段相关性重排 | query + candidates(≤10) → 排序后的 index/score |
接入 opencode 的配置示例(项目 .opencode/opencode.json):
opencode 插件在 integrations/opencode/jev-skill-router.ts:注册 jev_route_skill(task) 工具,请求进来时用 Jev 从本地 skill 目录选出最合适的一个并返回其完整指令。参考做法是配合 agent.build.tools.skill = false 关闭内置 skill 工具,使主模型上下文不再携带整个 skill 目录。
| 任务 | 数据 | 结果 |
|---|---|---|
| 间接注入检测 | InjecAgent 1,105 条 | 阈值 0.10:P/R 100%,良性误报 0% |
| 检索重排 | BEIR SciFact 900 对 | BM25 → Jev:MRR 0.622 → 0.843,Hit@1 50% → 78.3% |
| 意图分类 | SNIPS / Banking77 | 7 类 97.9% / 77 类 80.3% |
| 工具目录路由 | MetaTool 199 工具 | 相似干扰 k=5 96.5% |
| Skill router | SkillRetBench 501 库 | hybrid 架构 R@1 75.8%(最强基线 38.0%) |
| 命令风险门控 | 自建 130 条 | 危险拦截 100%、正常放行 98.2% |
| 模型难度路由 | RouterBench | 51%(无信号,负结果) |
| 轨迹失败归因 | Who&When 1,403 步 | AUROC 0.56(负结果) |
总计约 22,500 次 API 调用、52.2M input tokens、$2.19。
choice 最多 255 个选项、仅文本输入、state+questions 共享约 32k tokens、英语为主。jev-1.13.0,模型升级后建议重新评测。No reviews yet — be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/jev)<a href="https://allmcps.com/mcp/jev"><img src="https://allmcps.com/api/badge/jev?style=directory" alt="Jev on AllMCPs" /></a>