Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 30 tools.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
π‘ Paste into ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows)
Disclaimer: Community-maintained open-source project. Not affiliated with, endorsed by, or sponsored by the vLLM or Ray projects or any inference-serving vendor. Product and trademark names belong to their owners. MIT licensed.
Governed AI-ops for GPU inference clusters β vLLM (OpenAI API + Prometheus
/metrics) and Ray Serve / Ray Jobs (Ray dashboard), plus the single-process
serving engines SGLang and TGI (Text Generation Inference) β with a
built-in governance harness: unified audit log, policy engine, token/runaway
budget guard, undo-token recording, and descriptive risk-tier labels on every
audit row. It parses each engine's Prometheus /metrics directly (no Prometheus
server required) and
probes the Ray dashboard independently. A bearer token is optional (many
stacks run open).
Serving engines. vLLM is the flagship (full Ray Serve control plane: scale, drain, autoscale, LoRA, hot-swap). SGLang and TGI are supported for engine-agnostic observability β health, running-model identity, request-latency metrics, queue depth, and latency RCA β read from each engine's own endpoints and metric names. Being single-process servers, they have no Ray-shaped scale/drain API: those writes return a teaching error pointing you at a real horizontal-scale layer (Ray Serve / Kubernetes / a load balancer).
The flagship value is root-cause analysis, wrapped in guarded reads and writes:
diagnose_latency_spike (flagship RCA) β when TTFT/TPOT/e2e latency
climbs, it correlates queue depth (running vs waiting), KV-cache
pressure / preemptions, and prefix-cache locality into a ranked cause
plus the specific knob to turn (add replicas, raise max-num-seqs, fix
routing, enlarge KV cache). Every flag is a number, not a black-box verdict.diagnose_low_utilization β the inverse: idle GPUs, over-provisioned
replicas, or routing that strands a cache-warm replica β what to scale down./metrics endpoint directly; no
Prometheus/Grafana deployment needed.ray start --head).It delivers inference-cluster operations β reads and writes β accurately and efficiently, and records every one of them. It does not decide whether a write is allowed to happen. That is the agent's judgement, or the permission of the environment you connect it with: restrict the network path so it can only reach the read/metrics endpoints, or run the Ray dashboard without its job-submission API, and the writes fail at the server β the place that actually owns the permission.
So there is no read-only switch, no policy file, no approval gate to configure.
The one thing the tool guarantees is that nothing is silent: every call, over
MCP and over the CLI alike, lands an audit row in
~/.inference-aiops/audit.db, and destructive writes still capture their
before-state and record an inverse where one exists.
Each tool declares a
risk_level, kept in agreement with its[READ]/[WRITE]documentation tag by a test, and carried into the audit row as a descriptive tier β so a reviewer can see at a glance that a row was a high-risk scale-to-zero. It is a label, not a gate.
Running a smaller / local model? See agent-guardrails.md β it lists the guardrails this tool enforces for you (so you don't spend prompt budget restating them) and gives a ready-made system prompt for what's left.
| Group | Tools | Count | R/W (risk) |
|---|---|---|---|
| Metrics & RCA (vLLM) | request_metrics, queue_depth, kv_cache_stats, diagnose_latency_spike, diagnose_low_utilization | 5 | read |
| Engine-agnostic (vLLM / SGLang / TGI) | engine_health, engine_inventory, engine_request_metrics, engine_queue_depth, diagnose_engine_latency | 5 | read |
| Ray Serve (read) | serve_deployment_list, deployment_status, replica_list, autoscale_config_get | 4 | read |
| Ray Serve (write) | scale_replicas_up, scale_replicas_down, scale_to_zero, autoscale_config_update, drain_replica | 5 | write (med / high) |
| Models / vLLM | model_list, model_info, model_is_sleeping, lora_load, lora_unload | 5 | read + write (med) |
Sleep Mode / vLLM (needs VLLM_SERVER_DEV_MODE=1) | model_sleep, model_wake | 2 | write (high / med) |
| Ray cluster / jobs / GPU | ray_cluster_resources, ray_dashboard_status, ray_job_list, gpu_utilization, ray_job_cancel, replica_restart | 6 | read + write (med / high) |
| Deploy lifecycle | model_deploy, model_undeploy, deployment_redeploy, routing_policy_update | 4 | write (med / high) |
| Cost | cost_per_token | 1 | read |
The engine-agnostic group works against any supported engine (including vLLM); use it for SGLang/TGI targets or a uniform view across a mixed fleet. The Ray Serve / cluster / deploy write groups are vLLM-only (Ray control plane) β they teach-and-refuse on a SGLang/TGI target.
23 read, 16 write. High-risk writes (scale_replicas_down,
scale_to_zero, drain_replica, lora_unload, model_sleep,
replica_restart, model_undeploy, deployment_redeploy) all support
dry_run + double-confirm; reversible writes record an undo descriptor.
Sleep Mode requires a dev-mode server. vLLM registers
/sleep,/wake_upand/is_sleepingonly when started withVLLM_SERVER_DEV_MODE=1. Against any other server these three tools report that the route is absent and why, rather than failing vaguely. Sleep Mode suspends the same model; it does not swap base models β serving a different base model means restarting vLLM with a different--model.
Run as an MCP server (stdio) for the full 39-tool surface:
The CLI is a convenience subset (init, overview, serve β¦, metrics β¦,
secret β¦, doctor, mcp); the full 39 tools are exposed via the MCP server.
Every MCP tool passes through the bundled @governed_tool harness. It does not
decide whether a write is permitted β see What this tool does, and does not,
decide above β but it records every call:
~/.inference-aiops/audit.db
(relocatable via INFERENCE_AIOPS_HOME).risk_level; it is a label, not a gate. INFERENCE_AUDIT_APPROVED_BY
/ INFERENCE_AUDIT_RATIONALE are optional annotations recorded when set,
never required.Behaviour is exercised by the test suite against mocked vLLM /metrics, vLLM
OpenAI API, and Ray dashboard responses. ~80% of the tool self-tests on a
laptop β vLLM on a single GPU or CPU-mock plus a local one-node Ray head. It
has not been run against a live production cluster; see
docs/VERIFICATION.md for the live-verification
checklist.
Unverified against real hardware / topology:
/api/nodes),The fastest live check is inference-aiops doctor; the full checklist lives in
docs/VERIFICATION.md.
This is the GPU-inference member of the AIops-tools family (governed AI-ops with audit + budget + undo + risk tiers). If a vLLM or Ray capability you need is missing, or your stack speaks a dialect these tools don't yet handle β open an issue or a PR. Contributions welcome.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/inference-aiops)<a href="https://allmcps.com/mcp/inference-aiops"><img src="https://allmcps.com/api/badge/inference-aiops?style=directory" alt="Inference AIops on AllMCPs" /></a>