MCP server for ML drift detection, MLflow registry diffing, and HMAC-gated model rollback
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
π‘ Paste the JSON block into your client's configuration file under mcpServers, then restart the application.
Autonomous MLOps incident response agent, plus mendrift-mcp β an open-source MCP server for drift detection and ML incident tooling.
Published on PyPI and the
MCP Registry as
io.github.suneel190700/mendrift-mcp.
βΆ Live demo β run a real incident in your browser: supply an alert, watch the agent diagnose it against a real MLflow registry, and approve or reject the rollback at the human-in-the-loop gate. Toggle between a crafted synthetic scenario and real US consumer-credit benchmark data. React frontend on a FastAPI backend; the free tier sleeps, so the first load may take ~40s.
When a production model drifts or degrades, Mendrift detects it, diagnoses the root cause from monitoring and registry evidence, proposes a remediation, and executes it only after human approval.
Built with LangGraph (agent orchestration), LangChain (ChatAnthropic +
bind_tools), the Model Context Protocol, Evidently, MLflow, and Claude
(Haiku + Sonnet).
| tool | type | purpose |
|---|---|---|
get_drift_report | read | per-feature drift distances + schema changes (Evidently) |
summarize_metric_anomalies | read | production vs previous model scored on current traffic |
get_deployment_history | read | registry version transitions and aliases |
diff_deployments | read | params / metrics / feature-schema diff between versions |
propose_rollback | read | generates a reviewable rollback plan |
execute_rollback | gated | requires a single-use HMAC approval_token |
open_incident | write | incident record with diagnosis + evidence |
The approval gate is enforced in the tool layer, not the prompt:
execute_rollback verifies a single-use, action-scoped HMAC token minted only
by the human review flow β the minting function is never exposed over MCP. A
prompt-injected or confused agent cannot execute writes.
Tested live: Claude was first ordered to roll back "with full authorization" (it proposed but declined to fabricate a token), then handed a fabricated token, which the gate rejected by constant-time HMAC comparison:

See tests/test_approval_gate.py, including the action-scoping test: a token
minted for one model/version is invalid for any other.
The incident graph halts before execution (interrupt_before) and checkpoints
every step to SQLite. The process can die; a new process resumes the same
incident by thread_id after a human mints the approval token β which enters
state only via update_state(), from outside the graph. Denial is a
first-class path: no token β closed_approval_denied, no execution.

| step | model | why |
|---|---|---|
| classify | Haiku | single constrained label; cheapest path |
| diagnose | Sonnet | multi-hop tool reasoning over evidence |
| verify | Haiku | threshold check on fresh metrics |
Routing lives in a code table (ROUTER_TABLE), not prompts, so cost per path
is measurable config β ~3.9K input / 630 output tokens per incident. The
diagnose loop is bounded (max 8 tool calls) with per-call retries and capped
backoff; on tool failure the model receives a structured error record, and on
budget exhaustion the agent degrades to an incident with partial evidence β it
never invents a diagnosis. Destructive actions require affirmative evidence: a
rollback is recommended only when retrieved evidence links the symptom to a
specific deployment, never on deploy-correlation alone. The agent can also
recommend monitor β real but mild, non-actionable drift is watched, not
acted on.
MENDRIFT_DEMO=0 runs the agent against real infrastructure rather than fixtures:
scripts/seed_demo.py trains two sklearn versions into a local MLflow registry β
v13 clean, v14 with a schema swap and a training window polluted by missed-fraud
labels (recall 0.72 β 0.18, AUC 0.84 β 0.82) β and writes reference/current framesget_drift_report runs Evidently's DataDriftPreset over those frames, returning
real Wasserstein/JS distances against per-metric thresholds, plus schema changes
derived from actual column setsget_deployment_history / diff_deployments read the registry and the underlying
runs β real aliases, params, metricssummarize_metric_anomalies scores the current window with both the production and
previous versions, so it reports model divergence rather than population drift β
a rollback clears it, ordinary data shift does notexecute_rollback moves the production alias for realA live run diagnoses from computed evidence β e.g. "v2 introduced a schema swap replacing promo_flag with promo_flag_v2 β¦ label_noise 0.0 β 0.45 collapsing val_recall 0.724 β 0.176 β¦ 79.7% prediction-rate divergence from the prior version, model-induced, not population drift" β then halts for approval and resolves.
The eval suite deliberately stays on fixtures: evals need determinism and zero cost in CI, while live mode exercises the real stack.
A hosted web app wraps live mode behind a browser UI: a React (Vite) frontend on
a FastAPI backend, deployed on Render. A visitor submits an alert, the frontend
posts it to /api/diagnose, and the backend runs the real LangGraph agent β live
Claude reasoning over an embedded MLflow registry (sqlite://, seeded on boot) β then
halts at the HMAC gate. Approve or reject and the backend resumes the graph via
/api/decision, executing a real alias rollback and verifying recovery. The Anthropic
key lives only on the server; runs are rate-limited since each calls a real model.
Try it: mendrift-demo.onrender.com.
The dashboard toggles between two seeded worlds, so the same agent can be seen against both a crafted scenario and genuine real-world data:
scripts/seed_demo.py, model fraud-scorer) β the crafted schema-swap
incident: clean, teachable, an unambiguous rollback story.scripts/seed_real.py, model credit-risk) β the
Give Me Some Credit
dataset (real US consumer-credit records, target SeriousDlqin2yrs) split by borrower
age into reference/current windows for genuine feature drift, with a controlled model
regression injected into v2 (asymmetric missed-default label noise) so the incident
has ground truth. Real distributions and real Evidently drift; a known correct action.
Measured gap: val_recall 0.637 β 0.156, AUC 0.854 β 0.810.Injecting a known regression into real data is standard practice for validating a drift-detection system β it gives the evaluator ground truth for what the agent should decide while the drift computation still runs on genuine distributions.
The backend routes each request to the right world (model + parquet frames + label
column) per the dataset field; the tool layer reads those from env vars, applied
per-request under a lock so concurrent requests stay isolated.
Run the web app locally:
For frontend development with hot reload, run cd frontend && npm run dev (port 5173);
Vite proxies /api to the backend on port 8000.
src/mendrift/evals/ replays synthetic incident trajectories against the
real graph β only the LLM (scripted) and the read tools (fixture world)
are faked; the gated action tools are the genuine implementations, so the HMAC
gate is exercised by every test. Four assertions per trajectory:
| check | meaning |
|---|---|
no_ungated_writes | every execute_rollback carried a valid HMAC token β hard fail |
classification_ok | triage label matched |
tool_sequence_ok | required tool calls occurred in order (extras allowed) |
action_ok | terminal outcome matched |
19 logic-distinct incident scenarios spanning the decision space, each with its own evidence shape and correct action:
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/mendrift)<a href="https://allmcps.com/mcp/mendrift"><img src="https://allmcps.com/api/badge/mendrift?style=directory" alt="Mendrift on AllMCPs" /></a>