The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Rancher MCP listing page.
Operate Rancher-managed Kubernetes through any MCP client — discovery, generic resource access, and curated operator workflows, wrapped in an audit-logged, rate-limited, confirmation-guarded safety model.
Quick start · Tool surface · Architecture · Safety model · Compatibility · Development
Rancher is how real fleets run Kubernetes — and it speaks two APIs (the legacy
Norman /v3 plane and the modern Steve /v1 plane), varies by version, and wraps
every cluster behind its own proxy. Pointing a generic Kubernetes MCP server at it
misses everything Rancher-specific; pointing an agent at raw kubectl gives up
auditability, guardrails, and the management-plane view entirely.
MCP Rancher is built for that reality:
read_only: true and every mutation is refused at the config layer, before any
guard even has to fire.206 tools: 178 read-only · 28 writes · 5 destructive — counted from the
registry itself, not by hand. docs/tool-manifest.json
is generated from the live FastMCP registry (make tool-manifest) and a CI
gate fails the build if it ever drifts from the code. Per-tool descriptions,
safety annotations, and parameters all live there; the narrative registry with
slice tracking is docs/tool-catalog.md.
| Layer | What it does | Examples |
|---|---|---|
| Discovery & schema | Explore what any instance can do | rancher_server_version, rancher_norman_schema_list, rancher_capability_domain_list |
| Generic engine | CRUD + actions + links + watch on any resource, both planes | rancher_steve_resource_list, rancher_norman_resource_action_invoke, rancher_steve_resource_watch |
| Curated reads | Typed, shaped responses across ~25 domains | rancher_pods_list, rancher_deployments_list, rancher_longhorn_volumes_list, rancher_policy_reports_list |
| Curated writes | Guarded mutations | rancher_deployment_scale, rancher_deployment_restart, rancher_cron_job_suspend, rancher_node_cordon, rancher_secret_create |
| Operator rollups | One-call triage | rancher_cluster_health_check, rancher_find_failing_pods, rancher_find_stalled_rollouts, rancher_project_health_summary |
Domains covered: clusters & nodes · projects & namespaces · workloads · pods & services · storage · networking · config & secrets (values masked) · certificates (keys masked) · RBAC · auth & identity · apps & catalogs · logging pipeline · Prometheus monitoring · policy reports · CIS compliance · backup operator · etcd backups · Longhorn · Fleet · provisioning · settings & features · alerts & notifiers.
All 206 stay exposed by default — every tool schema is deferred behind
Claude Code's own search, so a small default would only help other hosts at
the good host's expense. A constrained host (a small local model, a tight
context budget) can opt into a smaller surface via RANCHER_TOOLSETS; see
Toolset profiles under Configuration.
Once published to PyPI, it's one line: uvx rancher-mcp.
Every tool takes an optional instance argument. Instances flagged
read_only: true refuse all mutations at the settings layer.
Three layers, deliberately separate: discovery tells you what an instance can
do, the generic engine can touch anything it exposes, and curated tools
make the common paths typed, shaped, and self-describing (every response carries
suggested_next_steps). Most curated tools are generated from YAML descriptors
(catalog/curated_tools/) with a drift gate — the editorial decisions live in
descriptors, not boilerplate.
Built for the day an agent is pointed at the cluster that pays your salary:
| Guard | Behavior |
|---|---|
| Read-only instances | read_only: true refuses every mutation for that instance, before tool logic runs |
| Destructive confirmation | Deletes require an explicit typed phrase (e.g. "delete steve namespace foo") — no phrase, no delete |
| Tool annotations | Every tool declares readOnlyHint / destructiveHint / idempotentHint, so clients can gate UX on them |
| Audit log | Every mutation emits a structured event="audit" record — tool, operation, plane, instance, resource, outcome. Argument names only; values never logged |
| Rate limiting | Token-bucket on writes (default 60/min) — a runaway loop can't machine-gun your API |
| Secret & key masking | Secret values and certificate private keys are structurally absent from curated responses (reveal is an explicit generic-tool opt-in) |
| Structured errors | Guard rejections return typed error_code envelopes agents can branch on — never raw strings |
| Primary target | Rancher 2.9.3 (production-validated) |
| Compatibility floor | Rancher 2.6.5 (kept green via capability detection) |
| API planes | Norman /v3 + Steve /v1 (+ per-cluster Kubernetes proxy) |
| Transport | stdio |
Capability detection bridges version differences at runtime — no version-pinned builds, no "works on my Rancher." Both targets are exercised by the same test suite, and read paths have been validated live against both a 2.6.5 lab and a 2.9.3 production fleet (validation report).
| Variable | Default | Purpose |
|---|---|---|
RANCHER_URL | — | Rancher server URL (single-instance mode) |
RANCHER_TOKEN | — | API token (token-xxxxx:yyyyyyyyy) |
RANCHER_VERIFY_SSL | true | TLS verification |
RANCHER_INSTANCES_JSON | — | Multi-instance config (see above) |
RANCHER_DEFAULT_INSTANCE | first defined | Instance used when a tool call names none |
RANCHER_MCP_SERVER_NAME | rancher-mcp | Server identity announced to clients |
RANCHER_MCP_SERVER_DESCRIPTION | built-in | Server description announced to clients |
RANCHER_MCP_WRITE_RATE_LIMIT_PER_MIN | 60 | Write rate limit (0 disables) |
RANCHER_TOOLSETS | all | Comma-separated toolset profile(s) exposed at startup — see below |
RANCHER_TOOLS | — | Comma-separated tool names force-included on top of the selected profile(s) |
RANCHER_EXCLUDE_TOOLS | — | Comma-separated tool names removed, applied last — always wins over RANCHER_TOOLS |
The default is all: every tool stays exposed. Claude Code, the primary
host, defers every tool schema behind its own search, so a small default
would only help other hosts at the cost of making the good host worse — this
is a deliberate choice, not an oversight.
RANCHER_TOOLSETS opts a constrained host (a small local model, or a context
budget) into a smaller surface. Values are either a family name — one per
src/rancher_mcp/tools/ module (storage, workloads, pods_services,
rbac, …; see rancher_mcp.toolsets.FAMILY_REGISTRARS for the full list) —
or the cross-family core profile: a ~32-tool triage/orientation set
(rancher_find_*, the health/summary rollups, core list/get pairs, and the
generic rancher_{steve,norman}_resource_{list,get} escape hatches). core
cuts the tools/list payload from 206 tools / ~401 KB to 32 tools / ~78 KB
(~80% smaller). Profiles compose: RANCHER_TOOLSETS=core,storage gives the
triage set plus everything storage-related.
An unknown profile name fails loudly at startup rather than silently
producing an empty or shrunken surface. Calling a real tool that exists but
isn't in the active profile returns a structured TOOLSET_NOT_ENABLED error
naming the tool, the toolset that would enable it, and the env var to set —
never a bare "unknown tool", which would make a disabled tool indistinguishable
from a typo.
Shipping and stable for read, triage, and guarded write operations. Honest ledger of what's beyond that:
Work is tracked to the tool level: docs/tool-catalog.md
(every tool has a row, every gap a slice ID) and ROADMAP.md.
9443; run it serially with make integration-current to avoid
overlapping Docker resource demand with the legacy lab.tests/fixtures/; respx pins the HTTP boundary in tests.catalog/curated_tools/*.yml
descriptors; make check-codegen fails on drift.Stack: Python 3.12 · FastMCP · httpx · Pydantic v2 · structlog · uv
See SECURITY.md for the threat model, token guidance, and how to report vulnerabilities.