The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Tachibot MCP listing page.
67 AI tools. 12 providers. One protocol.
Orchestrate Perplexity, Grok 4.6, GPT-5.6, Gemini, Qwen, Kimi K3, and MiniMax M3 from Claude Code, Claude Desktop, Cursor, or any MCP client.
Get Started · View Tools · Documentation
If TachiBot helps your workflow, a star goes a long way.
storm, re3, de_generic). The library had creative thinking (what_if, innovate, alt_view, creative_use) and zero creative writing: against a 7-stage pipeline, stage 4 (structure) and stage 5 (draft) were empty, and stage 6 (revise) fell through to reflexion — which optimises toward correct, while prose revision is cutting, rhythm and specificity. 85 → 88 techniques.blog_writer — researched long-form in one call, composing storm and re3 rather than carrying a private copy of their text. A copy would drift the moment either template was edited, and drift silently, because both paths still emit well-formed prose. research: true runs the persona interview first and feeds its 3 surprising answers into the four-pass write. 66 → 67 tools; enabled in full, balanced and research_power.preview_prompt_technique(technique="auto") had no route for writing intent. RECOMMEND_RULES carried no matcher, so "write a blog post" fell through to CORE_TECHNIQUES.slice(0, 3) — chain_of_note / astute_rag / spotlighting — while list_prompt_techniques advertised auto as the way in. Writing verbs now route to storm, re3, de_generic.blog_writer truncated silently at maxTokens: 8000. Its output is several full renderings of the same piece (re3's draft, restructure and line-edit passes, plus storm's 16 marked answers), unlike sibling tools that emit one artifact. 8000 was already at the ceiling at the word count the docs use as their own example, and nothing errored on truncation. Raised to 16000.data_assert states the downstream purpose first, then checks the data with executable assertions carrying counts, because a filter that silently matched zero rows is indistinguishable from one that worked. Plus chain_of_table, sub_table_first, sql_stages, xlsx_map, table_format, self_debug.astute_rag answers from memory first, then marks each source agrees/fills/conflicts and resolves every conflict out loud; crag grades sources correct/ambiguous/wrong and acts on the grade, since search always returns its nearest neighbours and has no way to return nothing; hyde drafts the answer you wish existed, searches with its vocabulary, then discards it.untrustedFence() called randomBytes(3) while its own comment block had reasoned its way to 4. Spotlighting's entire security property is that an attacker cannot guess the delimiter needed to close your block, so the shortfall was load-bearing. Now 8 hex characters, as documented.37 survived three expansions of the catalogue (37 → 74 → 85) inside the list_prompt_techniques description — the text a model reads to decide whether to call the tool at all. It is now derived from FLAT_TECHNIQUES.length.spotlighting fences it and states the rule. Few-shot (5): nothing touched example selection or ordering, despite this moving accuracy more than most reasoning tricks.hourglass (compress to a rule, discard the exploration, regenerate from the rule instead of appending errors) and backtrack (rewind to the last validated checkpoint rather than restarting or patching).RECOMMEND_RULES gained eight groups matched to how people actually phrase the situation. Without these the new techniques were reachable only by knowing their name in advance.openai_reason and openai_search advertised "GPT-5.2" while the code calls the current OpenAI flagship tier.grok_search finally says why to pick it. Its entire description was "Web search" — no way to choose it over the four other search tools. It now states its live X/news grounding edge and cross-references grok_search_lite as the cheaper path. The five OpenRouter reasoners (deepseek_reason, glm_reason, stepfun_reason, ernie_reason, qwen_reason) likewise differed only by vendor trivia; each now carries an actual routing rule for when to pick it.heavy_coding profile from the TACHIBOT_PROFILE help, and its OPENAI_API_KEY hint claimed that key powers the Qwen and QwQ tools (those route via OpenRouter).scripts/package-extension.sh ran npm install --production, pruning devDependencies and leaving tsc unable to build afterward. It now restores the full dependency tree when it finishes.serverInfo.version was hardcoded from a 12-tool era ~28 minor versions ago, so Claude Desktop's connector panel misidentified every install. It now reads the real version from package.json. The /setup wizard's profile sizes were stale in the same way and now match the six real profiles.grok_search_lite is not broken, and is staying. That release claimed grok-4.3 "does not invoke web search either" — which would have made the cheap search tier pointless. Re-probed Aug 15 asking today's date: grok-4.3 ran 2 web searches and returned a citation in 10.4s; grok-4.6 did the same in 36.4s. Lite is grounded, cheaper, and 3–4x faster — prefer it for high-volume lookups. grok-4.5 is still the one that doesn't ground. The cheap tier now has its own grounding test, since nothing previously covered it.grok_search was not searching. It ran on grok-4.5, which never invokes the web_search tool on xAI's Agent Tools API — so it answered from training data while still rendering a source footer and a "Search used up to N sources" cost line computed locally from max_search_results, not from real usage. Probed Aug 14: asked today's date, grok-4.5 replied "October 10, 2025" with zero web_search_call entries; grok-4.6 replied correctly with two search calls and a citation. grok_search now runs grok-4.6 (same $2/$6 and 500K context as 4.5). The regression test asserts grounding — a web_search_call and ≥1 annotation — never the answer text, because a plausible ungrounded answer is precisely what hid this.grok_search_lite never ran on grok-4-1-fast at $0.20/$0.50. xAI retired that id and silently served grok-4.3 — HTTP 200, the swap disclosed only in the response body's model field, so nothing threw and no fallback fired. The advertised "~10x cheaper" was really ~1.6x.console.log calls were corrupting the MCP protocol stream. stdout is the JSON-RPC channel on a stdio server; diagnostics now go to stderr, guarded by a test that also rejects console.info, console.debug and process.stdout.write.create_workflow silently destroyed existing workflow files — the existence check guarded the directory, not the file. It now refuses to clobber unless overwrite: true, and reports the real .tachibot/workflows/ path instead of .tachi/workflows/."undefined" into system prompts for any unrecognised approach/task, because the || fallback sat inside the index. Fixed across grok_reason, grok_code, kimi_thinking, qwen_reason, deepseek_reason, glm_reason, stepfun_reason, ernie_reason.qwen_algo step called QwQ-32B with a generic prompt instead of Qwen3.8-Max, and qwq_reason lost its 4-persona deliberation entirely. Both now delegate to the real tools. Separately, planner_maker's qwen_coder step sent a parameter the schema rejects, so every execution of it failed validation.RUN_LIVE_TESTS=1, so an ordinary npm test no longer makes a billed API call.qwen/qwen3.8-max) — Alibaba's new flagship, GA the same day, now powers qwen_algo, qwen_reason, and the qwen_reason juror. 1M context (up from 262K), multimodal (text+image+video in), and the first Qwen exposing configurable reasoning effort. $2/$6 per M.qwen3-235b-a22b-thinking-2507 was correct but shallow; qwen3.7-max cost 1.6x for twice the wall time; qwen3-max-thinking was rejected outright — it returns zero reasoning tokens.medium, deliberately. This model's default effort behaves like high: 302s and $0.09 on a single qwen_algo call. At medium the same call answers at equal depth in 18–48s for $0.006–0.018 — cheaper and 3.5x faster than the model it replaces, which took 169s and $0.036 for a shorter answer. low starts dropping alternatives and is not used.reasoning_effort pass-through for OpenRouter — callOpenRouter now forwards the parameter; OpenRouter drops it for models that don't list it, so the quota fallback chain (3.8 Max → 3.7 Max → 235B Thinking) stays safe. Qwen3.8/3.7-Max also join the 600s extended-timeout bucket.qwen_coder, qwen_competitive, testgen stay on Qwen3-Coder-Next — it is coding-specialized and ~16x cheaper ($0.12/$0.80); 3.8 Max is the reasoning tier, not the codegen tier.moonshotai/kimi-k3) now powers every Kimi tool, the kimi juror, and the Kimi seat on diff_review. 2.8T open-weight MoE — the largest open model shipped — with a 1M context (up from 262K), native multimodal input, and long-horizon agentic coding that beats Opus 4.8 and GPT-5.5 on coding/agent benchmarks. Note the price: $3/$15 per M, 4x K2.7-Code — K2.7-Code stays as the automatic fallback (K3 → K2.7-Code → K2.6).gemini-3.7-flash) is the new search/workhorse tier behind gemini_search — 1M context at $0.75/$3.75 intro pricing through Dec 31, 2026 (then $1.50/$7.50); gemini-3.6-flash stays as the fallback. Flash-Lite remains gemini-3.5-flash-lite ($0.30/$2.50).grok-4.3 fallback is now quota/region insurance rather than a rollout workaround.grok-4.3 while xAI's region-staged rollout completes (EU mid-July) — tools keep working everywhere, and 4.5 activates by itself.grok_search_lite (new tool, 65 total) — the same Grok live search on grok-4-1-fast ($0.20/$0.50, 2M ctx), ~10x cheaper than grok_search. Use it for high-volume lookups and jury/council fan-outs.openai_* tools move to gpt-5.6-sol (flagship, same $5/$30 as 5.5 but stronger), terra for code (5.5-level at half price), luna for explanations ($1/$6). The $30/$180 gpt-5.5-pro tier is replaced by sol + reasoning effort; a permission fallback (sol → terra → 5.5) covers org-gated accounts./test and /audit skills (19 skills total) — /test generates runnable tests via testgen; /audit runs an OWASP/CWE security review via security_review.tachibot init now offers to install Claude Code skills with a per-skill skip choice ([Enter]=all · [s]=choose which to skip · [n]=none). Skills are opt-in — postinstall no longer writes to ~/.claude silently (npm run install-skills still installs all non-interactively)..mcpb extension now points at a valid entry point (was broken) and tracks the package version; tachibot init exits cleanly on non-interactive/CI stdin instead of hanging.refine_prompt (new tool) — opt-in prompt improver on a cheap/fast model: raw query → goal-first brief + what changed + open questions. Never auto-fires, never executes anything — you review, then use the brief. In Claude Code, /prompt refine presents the open questions as clickable choices and merges your answers into a final brief.list_prompt_techniques now defaults to the ~9 core techniques that still help 2026 reasoning models (output contracts like scot, pre_mortem, bdd_spec); all=true for the full 31.technique="auto" — preview_prompt_technique recommends the right technique for your task, with reasons. Ask tachi "improve my prompt" for the symptom-based menu.tachibot init (new CLI wizard) — detects your API keys and clients, prints the exact config for Claude Code and Claude Desktop. Never writes or echoes keys..mcpb from the latest release and double-click. No JSON editing.doctor — shows which keys are set, which tools are visible vs hidden and why, and what to try first.debug_triage — ranked root-cause hypotheses with the cheapest discriminating check for each (Grok 4.3)spec_writer — loose request → reviewable spec: user stories, Given/When/Then, out-of-scope, open questions (GPT-5.5)diff_review / plan_critique / testgen / security_review — multi-model diff review, adversarial plan red-team, test generation, OWASP/CWE audit/review, /redteam, /spec, /triage, /setupfocus orchestration screen: 37 lines of repeated scaffolding → 10 focused linesnpm test exits 0 again (uncancelled race timers leaked past Jest teardown)TachiBot ships with 19 slash commands for Claude Code. These orchestrate the tools into powerful workflows:
| Skill | What it does | Example |
|---|---|---|
/setup | Guided configuration — runs doctor, walks through keys/profiles | /setup |
/spec | Request → reviewable spec before planning | /spec add OAuth somehow |
/blueprint | Multi-model planning → bite-sized TDD steps | /blueprint add OAuth with refresh tokens |
/judge | Multi-model council - parallel analysis with synthesis | /judge how to implement rate limiting |
/think | Sequential reasoning chain with any model | /think grok,gemini design a cache layer |
/focus | Mode-based reasoning (debate, research, analyze) | /focus architecture-debate Redis vs Pg |
/breakdown | Strategic decomposition with pre-mortem | /breakdown refactor payment module |
/decompose | Split into sub-problems, deep-dive each one | /decompose implement collaborative editor |
/prompt | Recommend the right thinking technique (37 available) | /prompt why do users churn |
/algo | Algorithm analysis with 4 specialized models (DeepSeek lead) | /algo optimize LRU cache O(1) |
/lens | Long-context analysis over Kimi's 1M window | /lens find inconsistencies in this spec |
/reflect | Grounded reflexion loop — critique vs external evidence | /reflect harden this auth middleware |
/tot | Tree-of-Thought: branch → jury-prune → synthesize | /tot design a rate limiter |
/review | Multi-model diff review — panel + Gemini judge verdict | /review (or paste a diff) |
/redteam | Adversarial plan red-team — pre-mortem, risks, plan edits | /redteam <paste plan> |
/triage | Ranked root-cause bug triage | /triage <paste stack trace> |
/test | Generate runnable tests (edge cases first) | /test src/auth.ts |
/audit | Security review — OWASP/CWE findings + fixes | /audit the login handler |
/tachi | Help - see available skills, tools, key status | /tachi |
Skills automatically adapt to your configured API keys. Even with just 1-2 providers, all skills work.
Getting started? Type
/tachito see what's available.
gemini-3.7-flash, GA Aug 13 2026) — Flash/search tier; reasoning default stays gemini-3.1-pro-preview (Google has still not shipped a 3.5 Pro)| Profile | Tools | Best For |
|---|---|---|
| Minimal | 14 | Quick tasks, low token budget |
| Research Power | 38 | Deep investigation, multi-source |
| Code Focus | 43 | Software development, SWE tasks |
| Balanced | 56 | General-purpose, mixed workflows |
| Heavy Coding | 59 | Max code tools + agentic workflows |
| Full (default) | 67 | Everything enabled |
Detects your keys and clients, then prints the exact config for Claude Code and Claude Desktop.
Then verify with /mcp. Add API keys with --env, e.g. --env OPENROUTER_API_KEY=sk-or-xxx --env PERPLEXITY_API_KEY=pplx-xxx.
One-click (easiest): download tachibot-mcp.mcpb from the latest release and double-click it — Claude Desktop installs the extension with no JSON editing. Add your API keys when prompted (or later via the extension settings).
Gateway Mode (Recommended) — 2 keys, all providers:
Direct Mode — One key per provider:
Get keys: OpenRouter | Perplexity
See Installation Guide for detailed instructions.
perplexity_ask · perplexity_reason · grok_search · grok_search_lite · openai_search · gemini_search
grok_reason · openai_reason · qwen_reason · qwq_reason · kimi_thinking · kimi_decompose · deepseek_reason · glm_reason · stepfun_reason · ernie_reason · planner_maker · planner_runner · list_plans · spec_writer · blog_writer
kimi_code · grok_code · grok_debug · qwen_coder · qwen_algo · qwen_competitive · deepseek_algo · minimax_code · minimax_agent · testgen · debug_triage
gemini_analyze_text · gemini_analyze_code · gemini_judge · jury · diff_review · plan_critique · gemini_brainstorm · openai_brainstorm · openai_code_review · openai_explain · grok_brainstorm · grok_architect · security_review · kimi_long_context
think · nextThought · focus · tachi · doctor · usage_stats
workflow · workflow_start · continue_workflow · list_workflows · create_workflow · visualize_workflow · workflow_status · validate_workflow · validate_workflow_file
list_prompt_techniques · preview_prompt_technique · execute_prompt_technique · refine_prompt
local_query — any OpenAI-compatible local server (Ollama / LM Studio / llama.cpp / vLLM). Zero-cost, offline, private; also available as the local jury juror (hermes is accepted as a legacy alias). Runs whatever LOCAL_LLM_MODEL points at — e.g. a Nous Hermes build (ollama pull hermes3). Note the Hermes agent itself is model-agnostic — it runs on 300+ backends (GPT, Claude, Gemini, DeepSeek, or self-hosted Ollama/vLLM) — so "Hermes" was never a guarantee of distinct weights.
Contributions welcome! See CONTRIBUTING.md for guidelines.