Persistent decision memory for agents β learns which action works in which context
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
We haven't yet run this listing's install command through our automated sandbox check. This isn't a red flag β we're steadily working through the catalog.
π‘ Paste the JSON block into your client's configuration file under mcpServers, then restart the application.
The official Python client and Model Context Protocol (MCP) server for BanditDB β the ultra-fast, lock-free Contextual Bandit database written in Rust.
BanditDB abstracts away the complex linear algebra of Reinforcement Learning (LinUCB, Thompson Sampling) behind a dead-simple API. Build real-time personalizers, dynamic A/B tests, and give LLM agents mathematically rigorous persistent memory.
Requires the BanditDB Rust server running (default: http://localhost:8080).
The client features automatic connection pooling, exponential backoff retries, and strict timeouts.
Health
| Method | Description |
|---|---|
health() | Returns True if the server is reachable and the WAL writer is healthy. |
health_detail() | Returns the full health dict including per-campaign entropy and status ("ok" / "warning" / "critical"). |
Campaigns
| Method | Description |
|---|---|
create_campaign(campaign_id, arms, feature_dim, alpha=1.0, algorithm="linucb", metadata=None) | Register a new campaign. algorithm accepts "linucb", "thompson_sampling", NeuralLinUCBConfig, or ProgressiveConfig. metadata is an arbitrary JSON dict (β€ 64 KB). |
list_campaigns() | Returns a list of all campaigns (active and archived) with alpha, arm_count, and algorithm. |
campaign_info(campaign_id) | Returns full per-arm state: theta, theta_norm, prediction and reward counters. Raises APIError (404) if not found. |
report(campaign_id) | Business-level convergence report. converged=True means one arm has a statistically significant lead at 95% CI β safe to stop. converged=False means leading but CIs still overlap. converged=None means not enough data yet (< 30 rewards per arm). |
diagnostics(campaign_id) | Operator diagnostics: per-arm theta norms, A_inv uncertainty bounds, entropy health (selection_entropy, entropy_status, entropy_trend, likely_cause, suggested_action), tournament traffic, and neural buffer size. |
archive_campaign(campaign_id) | Soft-delete: pauses predictions/rewards but preserves all learned weights. Recoverable with restore_campaign(). |
restore_campaign(campaign_id) | Restore an archived campaign to active status with all weights intact. |
delete_campaign(campaign_id) | Permanently delete a campaign. Returns False if not found. |
Predict & Reward
| Method | Description |
|---|---|
predict(campaign_id, context) | Returns (arm_id, interaction_id). Pass interaction_id to reward() to close the loop. |
batch_predict(predictions) | Predict for up to 100 campaign/context pairs in a single round-trip. Each item: {"campaign_id": str, "context": List[float]}. Returns list of {arm_id, interaction_id} or {error} per item. |
reward(interaction_id, reward) | Record outcome. reward must be in [0.0, 1.0]. Raises APIError if the interaction has already been rewarded or has expired (default TTL: 24 h). |
Data & Export
| Method | Description |
|---|---|
checkpoint() | Flush WAL, snapshot models, write Parquet shards, run neural retrain + tournament eval, rotate WAL. Returns a summary string. |
export() | List Parquet export shards grouped by campaign. Returns {export_dir, shards}. |
Standard LLM agents are stateless β if they route a task to the wrong model and fail, they repeat the same mistake tomorrow. BanditDB's built-in MCP server gives the entire agent swarm shared persistent memory.
Add to your Claude configuration file:
~/Library/Application Support/Claude/claude_desktop_config.json%APPDATA%\Claude\claude_desktop_config.jsonThe agent swarm now has nine tools:
| Tool | What it does |
|---|---|
create_campaign | Create a new decision campaign. Accepts algorithm ("linucb" or "thompson_sampling") and alpha. Use Thompson Sampling for natural Bayesian exploration with no tuning needed. |
list_campaigns | List all active campaigns (shows algorithm and alpha) β useful to check what exists before calling get_intuition. |
campaign_diagnostics | Inspect per-arm learning state: theta_norm, prediction counts, reward rates, and entropy health. Use when a campaign doesn't seem to be learning or one arm is dominating. |
campaign_report | Business-level convergence report. Tells you whether the campaign has statistically converged and which arm is winning with confidence intervals. |
get_intuition | Ask BanditDB which arm to pick for a given context. Returns the arm and an interaction_id to save. |
batch_get_intuition | Get decisions for multiple campaigns in a single round-trip. Pass a list of {campaign_id, context} dicts. |
record_outcome | Report whether the chosen action succeeded (1.0) or failed (0.0). Updates the shared model. |
archive_campaign | Soft-delete a campaign. Pauses predictions/rewards but preserves all learned weights. |
restore_campaign | Restore an archived campaign to active status with all weights intact. |
Every decision made by any agent in the network improves the routing for all future agents.
BanditDB event-sources every prediction and reward to a Write-Ahead Log (WAL). Calling checkpoint() compiles completed predictionβreward pairs into Snappy-compressed Parquet files β one per campaign β for offline analysis with Polars or Pandas.
Every prediction is guaranteed to appear in the Parquet file even if its reward arrives hours later: BanditDB re-emits in-flight interactions at each checkpoint so delayed rewards are always captured in a future cycle.
The SDK ships three OPE estimators in banditdb.eval. They answer the question: "what would my average reward have been under a different policy β without running a live experiment?"
Install the eval dependencies:
| Estimator | Function | How it works | When to use |
|---|---|---|---|
| Replay | replay(df) | Accepts each interaction with probability (1/K) / propensity (Li et al. 2010). Unbiased sample of the uniform random policy. | Sanity check baseline. Low coverage is expected β ~1/K of interactions are used. |
| IPS / SNIPS | ips(df, clip=10.0) | Uses every interaction with importance weight (1/K) / propensity. Self-normalised to reduce variance. Weight clipping (default 10Γ) controls the bias-variance tradeoff. | Primary estimator. Use when you have enough data but want full coverage. |
| Doubly Robust | doubly_robust(df, clip=10.0) | Fits a linear reward model, then applies an IPS correction on residuals. Consistent if either the reward model or the propensities are correct. | Best statistical efficiency. Use when comparing multiple policies or sweeping alpha. |
All three estimators:
ValueError for Thompson Sampling campaigns (propensity column is null β TS does not log propensities)OPEResult with estimate, std_error, n_used, n_total, and methodNo reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/banditdb)<a href="https://allmcps.com/mcp/banditdb"><img src="https://allmcps.com/api/badge/banditdb?style=directory" alt="BanditDB on AllMCPs" /></a>