The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Agent Pattern MCP listing page.
MCP server that provides AI agent pattern expertise: generate, analyze, and evaluate agent system designs against a curated catalog of 61 agent patterns (ReAct, supervisor-worker, reflexion, self-RAG, LLMCompiler, and more).
Ask your agent (or call the tool directly):
Use design_agent_system to design a research assistant that combines web search with sandboxed code execution for multi-hop questions. Domain: tool-use-tasks.
The tool runs the full pipeline — analyze (pattern retrieval + requirements-weighted scoring) → generate (LLM structured output) → evaluate (metric scoring) → refine (bounded retry loop) — and returns a complete AgentSystemDesign with agents, relationships, tool contracts, and quality scores.
List all agent patterns in the tool_use category.
Get the full JSON of the react pattern.
submit_agent_design_job + get_agent_design_statusONLY for clients with short request timeouts (Cursor, Claude Desktop, TS-SDK). The default is design_agent_system with heartbeat defence. submit_agent_design_job returns a job_id immediately; poll get_agent_design_status until done:
submit_agent_design_job returns a job_id in milliseconds. The pipeline runs in a background task. Poll get_agent_design_status(job_id) every 10–30 seconds. When status is completed, the full design is in the result field. Cancellation is best-effort — the job exits at the next pipeline stage boundary.
This is the only fix that works for TS-SDK clients (Claude Desktop, Cursor).
The job store is SQLite at ~/.config/agent-pattern-mcp/jobs.db (configurable via AGENT_PATTERN_JOBS_DB).
| Tool | Description |
|---|---|
design_agent_system | Full pipeline: analyze → generate → evaluate → refine. Returns complete design + evaluation + quality metrics. Long-running (5–10 min); use this unless your client has a short request timeout. |
analyze_agent_system | Analyse requirements and derive agent pattern recommendations using pattern matching and domain similarity. Long-running (LLM call). Not idempotent. |
generate_agent_system | Generate an agent system design from requirements, topology, domain, and selected patterns. Long-running (LLM call). Not idempotent. |
evaluate_agent_system | Evaluate an agent system design against specified criteria and domain using pattern benchmarking. Long-running (LLM call). Not idempotent. |
list_agent_patterns | List all 61 patterns; filter by category and/or domain |
get_agent_pattern | Get full JSON for a specific pattern by name |
submit_agent_design_job | Start a background design job and return a job_id immediately. ONLY for clients with short request timeouts (Cursor, Claude Desktop, TS-SDK). For other clients use design_agent_system. Poll get_agent_design_status every 10–30 s. |
get_agent_design_status | Poll job status. Returns the current status, progress message, and the full design output when completed. |
cancel_agent_design | Cancel a running job (best-effort; takes effect at the next pipeline stage boundary; may take up to one LLM call). |
The server also exposes four user-invoked workflow prompts (slash commands in clients that support them):
| Prompt | Args | What it does |
|---|---|---|
design_agent_system_workflow | requirements* | Full analyze → generate → evaluate pipeline |
explore_pattern_catalog | domain, category | Live catalog discovery with embedded pattern names |
evaluate_my_agent_system | focus | Guide evaluation criteria + finding prioritisation |
compare_agent_topologies | topology_a*, topology_b*, requirements* | Two designs side-by-side; ~2× token cost |
* = required argument
In tool-only clients, the prompts are also exposed as tools via FastMCP's PromptsAsTools transform — you can call them like any other tool.
AI coding agents (Claude Code, OpenCode, Codex CLI) can load a SKILL that teaches them how and when to use this server's tools — including timeout-aware entry-point selection, output interpretation, and the full workflow recipe.
The SKILL lives in skills/agent-pattern-mcp/:
For agents that support file-based skills (OpenCode, Claude Code): point the agent's skill loader at skills/agent-pattern-mcp/SKILL.md. The skill tells the agent:
requirements, domain, and topology as separate structured argumentsfinal_quality_score, attempts > 1, and evaluation.recommendationsdesign_agent_system directly61 agent patterns across 10 categories (reasoning, tool_use, planning, reflection, research_synthesis, multi_agent, memory, retrieval, safety_control, observability) and 8 topologies (single-agent-loop, hierarchical, pipeline, plan-execute, parallel-fan-out, evaluator-loop, graph-orchestrated, swarm).
Each pattern/*-pattern.json file contains: name, category, topology, context, benefits, tradeoffs, quality_attributes (7 dims, 1-10), suitable_domains, unsuitable_domains, use_cases, avoid_when, component_types, technology_stack, anti_patterns, migration_from, migration_to, design_principles, best_practices, references.
Pull olkowa/agent-pattern-mcp without building. The hub compose starts the
MCP server only; start the TEI sidecars (olkowa/pattern-tei-embed,
olkowa/pattern-tei-rerank) separately and wire them via EMBEDDER_BASE_URL
/ RERANKER_BASE_URL:
See config/config.json for the full annotated example. Key sections:
simple / reciprocal_rerank), reranker settings, quality thresholds, blend weights, topology score threshold.heartbeat_enabled and heartbeat_interval_seconds.*-pattern.json files are loaded from. Defaults to ~/.config/agent-pattern-mcp/pattern for local runs; the Docker image sets PATTERN_DIRECTORY=/app/pattern, the baked 61-pattern catalog.The generator LLM is accessed through the LlamaIndex LiteLLM integration (llama-index-llms-litellm). All provider settings therefore follow LiteLLM's model syntax: <provider>/<model> (e.g. openai/gpt-4o-mini, anthropic/claude-sonnet-4-5, openrouter/minimax/minimax-m2).
The server composes the LiteLLM model string from your configuration as generator.provider + generator.config.model:
| Config / env | Example | Resulting LiteLLM model string |
|---|---|---|
provider: "openai", model: "gpt-4o-mini" | GENERATOR_PROVIDER=openai, GENERATOR_MODEL=gpt-4o-mini | openai/gpt-4o-mini |
provider: "anthropic", model: "claude-sonnet-4-5" | GENERATOR_PROVIDER=anthropic, GENERATOR_MODEL=claude-sonnet-4-5 | anthropic/claude-sonnet-4-5 |
provider: "openrouter", model: "minimax/minimax-m2" | GENERATOR_PROVIDER=openrouter, GENERATOR_MODEL=minimax/minimax-m2 | openrouter/minimax/minimax-m2 |
If the configured model already contains a provider prefix (e.g. openai/gpt-4o-mini), that prefix is stripped and replaced by the configured provider.
<provider>/<model> syntax: LiteLLM Providers documentationGENERATOR_BASE_URL (generator.config.base_url) — it is passed as the LiteLLM api_baseGENERATOR_API_KEY is passed as the LiteLLM api_key; temperature, top_p, top_k, and stream map to the corresponding LiteLLM parametersA single LLM configuration — generator — serves all pipeline phases (planning, generation, reflection). Legacy PLANNER_* / REFLECTOR_* variables and config sections were removed; configs containing them are rejected with an error.
| Variable | Default | Purpose |
|---|---|---|
GENERATOR_API_KEY | (required) | LLM provider API key (passed to LiteLLM as api_key) |
GENERATOR_PROVIDER | openai | LiteLLM provider prefix: openai, anthropic, openrouter, … — see LiteLLM Providers |
GENERATOR_MODEL | gpt-4o-mini | Model name; final model string is <GENERATOR_PROVIDER>/<GENERATOR_MODEL> (LiteLLM syntax) |
GENERATOR_TEMPERATURE | 0.1 | LLM temperature (all phases) |
EMBEDDER_PROVIDER | tei | Embedder provider |
EMBEDDER_BASE_URL | http://127.0.0.1:8080 | TEI endpoint |
RETRIEVAL_MODE | reciprocal_rerank | Fusion mode |
RETRIEVAL_ENABLE_RERANKING | false | Enable TEI cross-encoder reranking |
RERANKER_BASE_URL | http://pattern-tei-rerank:8080 | Reranker endpoint |
TRANSPORT | streamable-http | stdio or streamable-http |
PORT | 8051 | HTTP port |
TASKS_HEARTBEAT_ENABLED | true | Enable heartbeat progress notifications during long tool calls |
TASKS_HEARTBEAT_INTERVAL_SECONDS | 30 | Heartbeat interval in seconds (keep below client idle timeout) |
AGENT_PATTERN_JOBS_DB | ~/.config/agent-pattern-mcp/jobs.db | SQLite path for async job state (job trio); ephemeral in Docker |
~/.config/agent-pattern-mcp/pattern/my-pattern-pattern.json:Set PATTERN_DIRECTORY or pattern_directory in config to point at the directory (or place files in the repo's pattern/ dir).
Restart the server. Invalid files are skipped with a warning (lenient loading); valid ones appear in list_agent_patterns immediately.
http://localhost:8061/mcp (note the /mcp path; dev default port is 8061, systemd uses 8051).docker compose logs agent-pattern-mcp for startup errors.0.0.0.0:8051 by default (in-container). If Docker is used, the dev compose publishes on host port 8061 ("${MCP_HOST_PORT:-8061}:8051"); systemd uses 8051.GENERATOR_API_KEY must be set (via .env, environment, or config).GENERATOR_PROVIDER=minimax, GENERATOR_MODEL=minimax/MiniMax-M2.7, GENERATOR_BASE_URL=https://api.minimax.io/v1.*-pattern.json and contain all required fields (name, context, category, topology, suitable_domains, quality_attributes with all 7 keys).Pattern validation failed warnings.jobs.db lives at ~/.config/agent-pattern-mcp/jobs.db inside the container; it is ephemeral (lost on restart). Restarting a container discards in-flight and completed jobs. To persist jobs across restarts, mount a volume and set AGENT_PATTERN_JOBS_DB to point at it.design_agent_system (and to a lesser extent analyze_agent_system, generate_agent_system, evaluate_agent_system) run multi-stage LLM pipelines that can take 5–10 minutes per call. This is inherent to the workload, not a bug: the generator LLM must process a large input payload — the selected pattern definitions from the 61-pattern catalog, your requirements, and the full output of every previous stage — and produce a large, strictly structured JSON document (agents, relationships, tool contracts, quality scores) one token at a time. The design_agent_system pipeline repeats generate → evaluate up to three times, so a single call can comprise 9+ LLM round trips.
MCP clients (AI coding agents, MCP SDKs) sit between the server and the LLM. Many implement a client-side idle timeout: if no data is received on the HTTP connection for some period (typically 30–120 seconds), the client aborts the request. The server is still working — the LLM is still generating — but the client closes the connection and reports a timeout error to the agent.
This is a client-side behaviour, not a server-side one. The server processes the full request correctly; the client simply gives up before the response arrives.
Affected clients (hardcoded short timeouts):
| Client | Timeout | Notes |
|---|---|---|
| Claude Desktop (TS-SDK) | 60 s | Hardcoded; does not reset on progress notifications |
| Cursor (TS-SDK) | 60 s | Same as Claude Desktop |
| Other TS-SDK based agents | varies | Most cap at 60–120 s |
These clients cannot be reconfigured to accept longer timeouts — the timeout is baked into the SDK.
Clients covered by the heartbeat defence:
| Client | Timeout | Defence |
|---|---|---|
| Claude Code | ~300 s | Heartbeat every 30 s resets idle timer |
| OpenCode | ~300 s | Heartbeat every 30 s resets idle timer |
| Codex CLI | ~300 s | Heartbeat every 30 s resets idle timer |
| Other HTTP-transport agents | varies | Most reset on any received data |
Works for these because their idle timers are reset by any incoming data — the heartbeat progress notifications sent from a parallel async task on the server are received by the client, resetting its clock.
Every long-running tool emits progress notifications from a parallel coroutine every 30 seconds (configurable via TASKS_HEARTBEAT_INTERVAL_SECONDS). As long as the client resets its idle timer on any received data, the request stays alive for the full duration of the pipeline.
TS-SDK clients (Claude Desktop, Cursor, etc.) do not reset their timeout on progress notifications.
For full control and compatibility with timeout-limited clients, three tools provide a durable job handle:
submit_agent_design_job returns a job_id in milliseconds. The pipeline runs in a background task. Poll get_agent_design_status(job_id) every 10–30 seconds. When status is completed, the full design is in the result field. Cancellation is best-effort — the job exits at the next pipeline stage boundary.
This is the only fix that works for TS-SDK clients (Claude Desktop, Cursor).
make clientThe example client in examples/agent_client.py is a direct Python HTTP client — it is not an MCP agent. It calls the server over HTTP without any MCP SDK, and therefore has no client-side idle timeout. It makes a single blocking request and waits for the full response, regardless of how long it takes.
make client is a development/demo tool. It demonstrates that the server correctly completes long requests — the timeout issue is purely a client-side problem. For production use with MCP agents, the heartbeat defence covers the majority of clients; the async job trio is the universal fallback.
The repo carries a layered verification program (see docs/verification.md and docs/testing-strategies.md):
| Layer | What it proves | Entry point |
|---|---|---|
| L1 | unit suite (behavioral oracle) + secret canary | make test-unit |
| L1b | MCP tool-surface boundary fuzz | make test-oracles |
| L2 | Hypothesis property oracles over the decision modules | make test-oracles |
| L3 | mutation testing (mutmut ratchet) + planted-bug garden | make test-mutations (nightly) |
| L4 | FizzBee exhaustive model checks + FG spec garden | make verify-fizz, make verify-fizz-garden |
| L6 | deterministic-simulation races over the jobs store | make test-oracles |
| L8 | property-ID ledger + NL-Doc cross-consistency | make verify-ledger, make verify-cross-consistency |
| L9 | structural invariants over LLM output (opt-in live pipeline) | uv run pytest tests/eval/ -m llm (AGENT_BENCH_LLM=1) |
| L10 | perf smoke canary (RUN_PERF=1) | nightly workflow |
Fast gates: make check-all (lint + types + dead code + deps) and
make test-all (unit + oracles). The nightly canary lives in
.github/workflows/verification.yml.
Images publish to Docker Hub (olkowa/*) and GHCR (ghcr.io/olk/*).
docker login) access.The docker-publish target also creates and pushes an annotated git tag v$(VERSION).
New GHCR packages default to private. Visit github.com/users/olk/packages, open each package → Package settings → Change visibility → Public.
See systemd/README.md for the full guide: file layout, install steps, day-to-day commands, updating, and uninstall. Summary:
Prerequisite: The shared TEI infra stack must be installed first — see
~/Projekte/Python/tei-infra/README.md. The systemd MCP stack joins thetei-sharedexternal network to reach the embedder/reranker.
MIT — see LICENSE.