Local research agent that verifies its own answers. Runs on Gemma 3 4B + Ollama, $0/query.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
π‘ Paste into ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows)
agentic-research-engine-oss
The best $0 research agent that runs on a laptop. Open-source end-to-end, reproducible, privacy-preserving. No cloud dependency by default; no telemetry; every LLM call, every source, and every verification decision is visible.
Local-first research agent that verifies its own answers. Runs on
Gemma 3 4B + Ollama (3.3 GB on disk) for $0/query; swaps to any
OpenAI-compatible endpoint with one env var.
| Interfaces | CLI Β· Textual TUI Β· FastAPI web GUI Β· MCP server (Claude Desktop / Cursor / Continue) |
| Pipeline | 8-node LangGraph (classify β plan β search β retrieve β fetch β compress β synthesize β verify); every node env-toggleable for ablation |
| Retrieval | SearXNG meta-search + trafilatura fetch + hybrid BM25 / dense / RRF; opt-in bge-reranker-v2-m3 cross-encoder |
| Reasoning | HyDE query expansion Β· FLARE active retrieval Β· Chain-of-Verification (Dhuliawala et al 2023) Β· ThinkPRM step critic |
| Domains | 6 presets (general Β· medical Β· papers Β· financial Β· stock_trading Β· personal_docs) β write your own in 10 lines of YAML |
| Plugins | load Claude plugins or agentskills.io skills from GitHub or local paths |
| Memory | opt-in local SQLite trajectory log with semantic retrieval; wipe anytime; no telemetry |
| Providers | OpenAI Β· Groq Β· vLLM Β· SGLang Β· Together Β· Ollama β any OpenAI-compatible endpoint via OPENAI_BASE_URL |
| Quality | 137 mocked tests, zero-network Β· honest live benchmarks published in RESULTS.md Β· MIT end-to-end |
| you currently use | we give you |
|---|---|
| Perplexity / ChatGPT Deep Research / Kagi Assistant | the same reasoning-with-citations flow, local and free, with your data never leaving the machine |
| Perplexica self-hosted | the UX Perplexica has plus a CoVe verifier, FLARE active retrieval, adaptive compute router, and Claude-plugin packaging |
| Khoj | stronger research-specific reasoning (we're not personal-knowledge-focused), six domain presets, and an MCP server for other agents to call |
| gpt-researcher | newer pipeline architecture, better small-model handling, observable trace, plugin ecosystem |
| MiroThinker-H1 / OpenResearcher-30B | they're stronger on BrowseComp; we run on a laptop with no GPU and cost $0 |
| Writing your own LangGraph research agent | save 2-3 months; reuse our 8-node pipeline + 30+ tested env gates + 137 tests |
Honest read: on complex multi-hop reasoning benchmarks, Gemma 3 4B sits 15β25% below 30 B+ open models. We don't claim to beat GPT-5.4 Pro. We claim to be the best $0, runs-on-your-laptop, fully-open research agent in April 2026.
Expected wall-clock on an M-series Mac: ~45 s for a factoid, ~90 s for multi-hop synthesis. Zero dollars per query.
Gemma 3 4B is surprisingly good at structure (plan, route, verify,
compress) but confabulates specific factoids when SearXNG doesn't
surface a source containing the right token. Live SimpleQA-mini run on
2026-04-21 (see engine/benchmarks/RESULTS.md)
showed gemma3:4b emitting "2023" for "year Anthropic published
Contextual Retrieval" (gold: 2024) and "LayoutLMv3" for "which
cross-encoder for reranking" (gold: bge-reranker-v2-m3).
The fix you probably want isn't a smarter synthesizer β it's a
more honest one. A 5-question head-to-head on the same retrieval
output showed gpt-5-nano + gpt-5-mini refuse to confabulate when
evidence was missing ("The provided evidence does not answer this
question"), where gemma3:4b confidently guessed. Per-claim
faithfulness went from 82.9 % β 100 %. Pass rate barely moved (1/5
vs 0/5) because retrieval is the real bottleneck β if SearXNG
didn't return a source with the gold token, neither model can
produce it.
Swap the whole stack to a cloud endpoint:
Cost is dominated by synthesizer tokens (~5β15 k per query). Full
cloud mode with gpt-5-nano + gpt-5-mini runs roughly
$0.02β0.05 per research query and is ~2-3Γ slower than Gemma
local (measured: 127 s vs 52 s mean wall on the 5-question subset).
Works with any OpenAI-compatible endpoint β Groq, Together, Mistral,
DeepSeek, local vLLM β so you can pick a cheap fast model
(llama-3.3-70b on Groq β $0.003/query) or a frontier one. Per-node
base-URL routing (run gemma3:4b locally for plan/verify AND gpt-5-mini
on cloud for synth in the same query) is tracked for 0.2; today the
pipeline uses one global OPENAI_BASE_URL.
The bigger accuracy lever is retrieval. Point
LOCAL_CORPUS_PATH at an indexed corpus containing your answer and
either model will be correct.
Five runnable notebooks in tutorials/:
Each notebook is self-contained, runs end-to-end on Colab free tier, no credit card required.
Three panes: sources Β· answer + hallucination flags Β· trace + memory hits. Press Enter to ask, Ctrl-M to cycle memory mode, Ctrl-L to clear, Ctrl-Q to quit.
localhost:8080)No auth. No cloud. No analytics. Dark theme. Streams tokens in place.
engine/ β the flagship8-node LangGraph pipeline with 2026-SOTA composition:
classify β plan β search β retrieve β fetch_url β compress β synthesize β verify
Every stage is env-toggleable for leave-one-out ablation. Techniques
folded in: HyDE, CoVe verification, iterative retrieval, FLARE active
retrieval, question classifier router, step critic (ThinkPRM pattern),
LongLLMLingua-lite compression, cross-encoder rerank
(BAAI/bge-reranker-v2-m3), Anthropic contextual chunking, W6 small-
model hardening (three-case synthesize prompt + per-chunk char cap).
core/rag/ β reusable retrieval primitives (v1 stable)HybridRetriever (BM25 + dense + RRF) Β· CrossEncoderReranker Β·
contextualize_chunks (Anthropic pattern) Β· CorpusIndex (bring-
your-own-PDFs). 5 exports, used by the engine and the archived
recipes.
archive/recipes/ β pre-engine reference recipesresearch-assistant, trading-copilot, document-qa,
rust-mcp-search-tool. All still work; all tests still pass. The
research-assistant/production/main.py is a thin shim over
engine.core.pipeline so the cookbook framing is preserved.
Six YAML files in engine/domains/:
| preset | when to use |
|---|---|
general | default; anything |
medical | disease / treatment / drug / trial (PubMed / Cochrane / NEJM bias; no prescriptive advice) |
papers | academic CS / ML / physics / biology (arXiv + Semantic Scholar + OpenReview) |
financial | SEC filings, earnings, company fundamentals (dates on every number) |
stock_trading | technical + news per ticker β hard rule: never recommends buy/sell/hold |
personal_docs | Q&A over your own corpus, air-gapped (only corpus:// URLs allowed) |
Write your own in ~10 lines of YAML β see docs/domains.md.
Supported formats: PDF (via pypdf), Markdown, plain text, HTML (via
trafilatura). The index persists as a directory with a human-readable
manifest.json + a pickled index.pkl. Rebuild anytime the docs change.
Details: docs/self-learning.md covers the
trajectory + memory model; docs/plugins-skills.md
covers external plugins.
engine/mcp/server.py is a Python MCP server exposing:
research(question, domain?, memory?) β structured {answer, verified_claims, unverified_claims, sources, trace, totals, memory_hits}reset_memory()memory_count()Bundled Claude plugin at engine/mcp/claude_plugin/ β four skills
(/research, /cite-sources, /verify-claim, /set-domain), ready to
submit to the Anthropic marketplace.
Register in Claude Desktop:
Install third-party Claude plugins or Hermes (agentskills.io) skills:
Safety: every install runs a forbidden-symbols scan
(eval(, exec(, os.system(, β¦) β rejects plugins that would
execute arbitrary code. Registry lives at
~/.agentic-research/plugins/, fully inspectable, wipable.
Full docs: docs/plugins-skills.md.
Every stage has an ENABLE_* flag so you can leave-one-out ablate.
Deep spec: docs/architecture.md.
Full list in engine/core/pipeline.py header. Most-common knobs:
| var | default | purpose |
|---|---|---|
OPENAI_BASE_URL | unset (cloud OpenAI) | route to Ollama / vLLM / Groq / etc. |
OPENAI_API_KEY | ollama | sentinel for local; real key for cloud |
MODEL_SYNTHESIZER | gpt-5-mini (cloud) or gemma3:4b (Mac-local path) | final-answer model. Swap to gpt-5, claude-sonnet-4-5, llama-3.3-70b on Groq, etc., for higher factoid accuracy while keeping the rest of the pipeline local. |
TOP_K_EVIDENCE | auto (5 for small, 8 for large models) | retrieval budget |
ENABLE_RERANK | 0 | opt-in; first run downloads bge-reranker-v2-m3 (~560 MB) |
ENABLE_FETCH | 1 | trafilatura full-page fetch |
ENABLE_STREAM | 1 | stream synthesis tokens to stdout |
ENABLE_TRACE | 1 | per-call observability + summary at CLI end |
LOCAL_CORPUS_PATH | unset | set to an index dir to augment search with your docs |
MEMORY_DB_PATH | ~/.agentic-research/memory.db | SQLite trajectory store |
Full list: docs/architecture.md env-vars section.
All tests are mocked β no network, no API key, no model downloads. Live
integration smokes are separate (make smoke).
CI runs on every push / PR touching engine / core / recipes β see
.github/workflows/engine-tests.yml.
| symptom | likely cause | fix |
|---|---|---|
ModuleNotFoundError: No module named 'engine' | PYTHONPATH missing the repo root | export PYTHONPATH=$(pwd) from the repo root |
| CLI answer is empty + fast | Ollama not running | ollama serve in another terminal, or ollama list to check |
Connection refused on :8888 | SearXNG not up | cd scripts/searxng && docker compose up -d |
Connection refused on :11434 | Ollama not running | ollama serve, or let the system service start it |
First make smoke hangs ~20 s before output | Model warming up on first request | normal; subsequent queries are faster |
ENABLE_RERANK=1 stalls on first run | 560 MB bge-reranker download | wait it out once; cached after |
[corpus] LOAD BROKEN | corrupt or wrong-version index | delete + rebuild via scripts/index_corpus.py |
| TUI shows gibberish over SSH | terminal too narrow | resize to β₯ 100 cols; Textual needs space for the 3-pane layout |
Web GUI shows Invalid memory mode | malformed POST | use the form UI; values validated against off/session/persistent |
| Streaming cuts off mid-answer | flaky backend | re-run; batched fallback kicks in on next attempt. Set ENABLE_STREAM=0 if it persists |
zsh: command not found: twine (or similar) after uv pip install <pkg> | uv's venv isn't auto-activated by your shell | use .venv/bin/<cmd> β¦, uv run <cmd> β¦, or source .venv/bin/activate before running |
bad interpreter: .../python3: no such file or directory after moving or renaming the repo dir | venv shebangs are absolute paths tied to the dir the venv was created in | recreate: rm -rf .venv && uv venv && uv pip install -e . (or re-install whatever you had) |
make test says 0 tests collected | wrong CWD | run from the engine/ dir or set PYTHONPATH |
| Claude Desktop doesn't see the plugin | plugin.json in wrong path | /plugin marketplace add <absolute-path-to>/engine/mcp/claude_plugin |
Still stuck? Open an issue with the bug_report
template β include ollama list, engine version, and the error.
engine/benchmarks/RESULTS.md β
verified_ratio 85.5 %, zero must_not_contain hits; the model
isn't emitting banned strings, it's picking wrong ones). Mitigations:
(a) swap the whole stack to a cloud endpoint (see "Higher factoid
accuracy" above β $0.02β0.05/query with gpt-5-nano + gpt-5-mini),
(b) give the engine a LOCAL_CORPUS_PATH so your own docs become
retrieval targets, (c) set ENABLE_RERANK=1 to bias retrieval
toward the right sources.CHANGELOG.md.tools_enabled field in presets finally activates), first LoRA run if GPU arrives, plugin catalog in docs/.Good first issues: CONTRIBUTING.md. RFCs for
anything pipeline-scope. Plugin + domain-preset submissions welcome.
No Co-Authored-By trailers; author-as-written-by.
MIT. See LICENSE.
agentskills.io skill format we interoperate with)This PyPI package is the official source of the MCP server registered at https://registry.modelcontextprotocol.io. The line below is the ownership marker the registry validates β do not remove when editing this README.
mcp-name: io.github.TheAiSingularity/agentic-research
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/agentic-research)<a href="https://allmcps.com/mcp/agentic-research"><img src="https://allmcps.com/api/badge/agentic-research?style=directory" alt="Agentic Research on AllMCPs" /></a>