Tells you whether a trading strategy's edge is distinguishable from luck.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
One-click editor setup isnβt available for this listing yet β we donβt have a confirmed install command, and weβd rather show nothing than point your editor at the wrong package or host. Follow the projectβs own setup instructions, linked above.
A Sharpe ratio is an estimate. Regimen tells you whether it is a fact.
Regimen takes a trading strategy's equity curve and answers two questions an aggregate performance number cannot: is the measured edge distinguishable from luck, and in which market conditions does it actually hold?
It is built for agents as much as for people. Everything is available over a REST API and over MCP, and the most common answer it gives is that the evidence is too thin to support the claim being made. That is the product, not a failure mode.
https://regimen-nu.vercel.app/mcp (revision 2026-07-28)io.github.RaYYeR220/regimen (listing)verification/README.md Β· Claims ledger β CLAIMS.md Β· Real vs simulated β MOCKS.md Β· Scorecard β EVAL.mdA strategy publishes a Sharpe ratio of 2.4 over six weeks and a 58% win rate over 40 trades. Both numbers are real. Neither is evidence.
A Sharpe ratio computed from a short, skewed, fat-tailed sample carries an error bar wide enough to swallow the claim. Forty trades cannot distinguish a 58% edge from a coin. And once a strategy has been re-tuned twenty times, the best configuration looks good for the same reason the tallest of twenty random people is tall.
This is not a niche statistical objection β it is the single most common way capital is lost to a backtest. The mathematics for handling it has existed since 2012 and is almost never applied, because it requires more than dividing a mean by a standard deviation.
Here is real output from the live service, on a curve with an annualised Sharpe of 3.72:
A dashboard would have printed 3.72 and stopped.
1. Significance. The Probabilistic Sharpe Ratio β the probability the true Sharpe exceeds a benchmark given the sample's length, skewness and kurtosis. The Minimum Track Record Length β how many periods would be needed before the claim could be made at all. The Deflated Sharpe Ratio β the same statement corrected for how many configurations were tried first. A stationary-bootstrap confidence interval that preserves serial dependence.
2. Regime attribution. Each period's return is joined to the market conditions that held on that UTC date β volatility, funding, open interest, positioning, sentiment, trend state β read point-in-time, so nothing in a bucket could only have been known afterwards. Every factor carries a permutation test: the observed dispersion of performance across buckets is compared against the dispersion produced by randomly reshuffling the regime labels, because slicing a return series eight ways guarantees a flattering subset. Without that p-value a regime map is a data-mining machine.
3. Self-attack. Every analysis can be run against controls whose answer is known in advance: the strategy's own returns with the mean removed (true Sharpe exactly zero, so a correct engine must grade it near 50%), and a simulated population of edgeless strategies matched for length and volatility, so the real result can be placed as a percentile against pure luck. The result is published, including when a control fails.
4. Divergence check. Where a source publishes its own figures, Regimen recomputes them from the equity curve and reports the difference. Differing conventions explain most gaps, but a user quoting a dashboard deserves to know when the curve underneath says otherwise.
That curve is deliberately too short, and Regimen says so rather than producing a number.
verification/README.md has a full-length example that
produces a graded verdict, plus the health and deployment-proof checks.
| Method | Path | What it answers |
|---|---|---|
POST | /api/v1/evaluate | Is this track record distinguishable from luck? |
POST | /api/v1/regime-map | Which market conditions is the edge concentrated in? |
POST | /api/v1/self-attack | Why should I believe the verdict? |
GET | /api/v1/status | Upstream reachability, cache occupancy, demo-key availability. |
GET | /api/v1/openapi.json | The machine-readable contract. |
GET | /api/health | Liveness and the exact build commit. |
Every response is { data, meta } or { error, meta }, where meta carries a request id,
the build commit, and which credential mode served the request. Errors carry a stable
machine-readable code, a retryable flag, and a remedy written to be actionable by an
agent rather than a human reading a stack trace.
Connect any MCP client to https://regimen-nu.vercel.app/mcp over Streamable HTTP. No
authentication is needed for the inline source. For a client configured by file:
To inspect it interactively: npx @modelcontextprotocol/inspector and point it at the
same URL.
Tools β regimen_evaluate_track_record, regimen_regime_map, regimen_self_attack,
regimen_describe_factors. Each advertises an outputSchema and returns validated
structuredContent; each is annotated readOnlyHint because nothing here writes, trades
or signs; each takes a detail switch so an agent can ask for the verdict and its reasons
rather than every bucket.
Resources β regimen://methodology (the statistics, in full), regimen://evidence-tiers
(the exact grading thresholds), and the template regimen://factor/{key}, whose key
argument supports completion/complete.
Prompt β validate_strategy, the full review in the right order, with instructions not
to lead with the annualised Sharpe.
The engine is written against a domain model β a track record, a regime series β and never against a vendor's response shape. A data source is a thin adapter that produces those two things. That is why the same analysis serves an OlaXBT Nexus strategy and a curve pasted in from a spreadsheet, and why adding a venue is an adapter rather than a rewrite.
OlaXBT Nexus is the live data source. The adapter reads the strategy's equity curve,
trades and published metrics, and reads eight market-condition factors per date with an
explicit as_of, which is what makes the regime attribution free of lookahead. Rate
limiting and caching live in the client, not at call sites: a Builder-tier key allows 80
requests a minute and a regime map wants hundreds of point-in-time reads, so calls are paced
under a token bucket and every immutable past-dated read is cached.
The statistics core carries 370 tests. Known-answer cases for PSR, MinTRL, the Deflated Sharpe Ratio, skewness, kurtosis, Wilson intervals and drawdown were generated independently of this implementation and carry their arithmetic in a comment. A negative-control test across 20 seeds confirms a zero-mean series clears 95% confidence on 2 of 20 runs β the nominal size β while the matching positive-mean series clears it on 20 of 20.
EVAL.md is a pre-registered graded evaluation: 82 synthetic strategies with known ground
truth, scored on false-positive rate, power, correct refusals on degenerate input,
calibration, and regime detection. The suite and its targets were fixed before the engine
was ever run against them.
It currently fails three of its six targets, and the failures are published rather than tuned away. They are worth reading, because they are the honest limits of the method:
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/regimen)<a href="https://allmcps.com/mcp/regimen"><img src="https://allmcps.com/api/badge/regimen?style=directory" alt="Regimen on AllMCPs" /></a>