Trace AI agent execution: every tool call, every error, every dollar. Open source, local-first.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
π‘ Paste into ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows)
Open source agent observability β see what your agents did, why they failed, and what it cost. Runs locally. No cloud required.
One command. No manual config. No copy-paste.
This auto-detects your AI agent (opencode, Claude Code, Cursor) and configures everything:
Then restart your agent. Every action will now self-report.
After running a task in your agent, the dashboard shows:
setup can't detect your agent)opencode: Add to opencode.json:
Claude Code: Create .mcp.json:
Cursor: Add to Cursor Settings > MCP:
See Agent-Specific Setup below for detailed per-platform instructions including self-reporting directives.
Note: All data is written to a local SQLite database in
~/.agent-observability/. No data leaves your machine.
For agents that can't self-report, wrap any MCP server and every tool invocation gets traced automatically:
Proxy mode captures only MCP tool calls (~30% of typical agent actions). Prefer the setup / MCP server approach above.
.mcp.json in your project root:.claude/instructions.md (or reference the existing SKILL.md under .claude/skills/agent-obs/SKILL.md):agent-obs dashboard, open http://localhost:9400, and look for your session after the agent completes a task.~/.cursor/mcp.json, or Settings β MCP β Add new global MCP server):.cursorrules (or .cursor/rules/agent-obs.md):agent-obs dashboard, open http://localhost:9400, and look for your session after the agent completes a task.| Method | Path | Description |
|---|---|---|
GET | /api/sessions | List all sessions. Query params: ?limit=20&offset=0&grade=B |
GET | /api/sessions/:id | Get full session detail with all tool calls, tokens, decisions, and grades |
GET | /api/sessions/:id/tool-calls | List tool calls for a session. Query params: ?status=error&server=filesystem |
GET | /api/sessions/:id/tokens | Get token usage history for a session |
GET | /api/sessions/:id/decisions | Get decision points for a session |
GET | /api/search | Full-text search across tool call inputs/outputs. Query param: ?q=read_file |
GET | /api/stats | Aggregate statistics: total sessions, avg grade, total tokens, total cost |
POST | /api/export | Export session data as JSON. Body: { "sessionIds": ["abc123"], "format": "json" } |
GET | /api/health | Health check. Returns { "status": "ok", "dbSize": "2.4MB", "sessionCount": 47 } |
| Tool | Parameters | Returns |
|---|---|---|
obs_record_tool_call | sessionId, toolName, serverName, duration (ms), status (success/error), input, output | { "id": "call-uuid", "recorded": true } |
obs_record_token_usage | sessionId, inputTokens, outputTokens, model | { "totalTokens": 1850, "estimatedCost": "$0.023" } |
obs_record_decision | sessionId, context, options, chosen, reasoning | { "id": "decision-uuid", "recorded": true } |
obs_record_grade | sessionId, grade (A/B/C/D/F), reasoning | { "grade": "B", "recorded": true } |
obs_get_session_report | sessionId | Full session JSON with all calls, tokens, decisions, grade |
obs_list_sessions | limit, offset | Array of { id, description, grade, createdAt, tokenTotal, estimatedCost } |
read_file, execute_command, search_codesuccess or errorCosts are estimated based on published API pricing:
| Model | Input (per 1M tokens) | Output (per 1M tokens) |
|---|---|---|
| Claude 3.5 Sonnet | $3.00 | $15.00 |
| Claude 3 Opus | $15.00 | $75.00 |
| GPT-4o | $2.50 | $10.00 |
| GPT-4 Turbo | $10.00 | $30.00 |
Costs are tracked per session and displayed in the dashboard. You can add custom model pricing via ~/.agent-observability/models.json.
Every session receives a letter grade based on efficiency and correctness:
| Grade | Label | Criteria |
|---|---|---|
| A | Clean | Zero errors, minimal token waste, no unnecessary tool calls |
| B | Minor issues | Some inefficiencies, but no failures |
| C | Inefficient | Excessive token usage, redundant tool calls, recoverable errors |
| D | Risky | Significant problems β failed tool calls, high cost, wrong tools chosen |
| F | Failed | Errors prevented task completion, or agent abandoned the session |
When the agent has multiple possible actions, it can record why it chose one over another. Example:
Decision points create an audit trail of agent reasoning, making it possible to understand not just what happened, but why.
Every action is timestamped and linked to a session. The audit trail answers:
Grades are assigned manually by the agent via obs_record_grade or automatically by the dashboard based on session statistics:
Grades accumulate over time. The /api/stats endpoint shows your average grade and grade distribution.
AI agents are a black box. You tell Claude Code or Cursor to fix a bug, refactor a module, or add a feature β and minutes later you have a diff. But what actually happened in between? How many tool calls did it make? How many tokens did it burn? Which files did it read? Did it try three approaches before landing on one? You have no idea.
This isn't just academic curiosity. Without observability:
Agent Observability gives you a flight recorder for every agent session. It answers: what happened, why it happened, and what it cost.
This project combines ideas from three prior projects:
AgentShelf scans a Shopify store's theme, apps, admin settings, and custom code to produce an AI readiness score. As it scans, it builds a full audit trail β every file read, every API call made, every finding logged with a timestamp and data source. Pattern adopted here: structured audit trail with per-action timestamps, data source attribution, and session-level summarization.
Veros generates synthetic longitudinal patient records (FHIR R4) and uses them to validate clinical decision support agents. Its core principle: "no trace, no answer." Every snippet of clinical reasoning produced by an AI must be traceable back to the specific data element (lab result, medication, condition) that supports it. Pattern adopted here: decision point tracking β every agent choice must cite its basis, making reasoning auditable and contestable.
MCP Observatory secures MCP servers β testing them for vulnerabilities, schema drift, and attack surfaces before agents depend on them. It's used by 850+ developers weekly and powers CI pipelines for MCP server security. Use Observatory to secure your MCP servers. Use agent-obs to trace the agents that depend on them. Observatory validates; agent-obs observes.
Pattern adopted here: tool health scoring, session grading (A-F), and automatic degradation flags.
This open source package (agent-observability on npm) is the core engine. It's MIT licensed and will always remain free and local-first:
A cloud version (app.agentobservability.dev) is in development with additional features:
The open source package will always be able to run independently. The cloud version is additive, not a replacement.
MIT
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/agent-observability)<a href="https://allmcps.com/mcp/agent-observability"><img src="https://allmcps.com/api/badge/agent-observability?style=directory" alt="Agent Observability on AllMCPs" /></a>