Plain-English training-data scout: 12 experts sweep HuggingFace + GitHub, compose scored packs.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
One-click editor setup isnβt available for this listing yet β we donβt have a confirmed install command, and weβd rather show nothing than point your editor at the wrong package or host. Follow the projectβs own setup instructions, linked above.
Plain-English topic in β 12 experts fan out over Hugging Face + GitHub in parallel β coverage-scored training pack JSON out in ~30 seconds.
No API keys. No Bridge. No account. Works out of the box once connected to any MCP host β no LLM inside; your agent brings the thinking.
Harvest is a training-data scout: you describe what you want to train on in plain English
("image classification", "medical question answering"), and 12 domain experts β
code, math, science, language, vision, audio, medical, law, knowledge, safety, affect, systems β
each sweep Hugging Face datasets and GitHub repos with their own curated queries. The local mixer
then composes the hits into a balanced, coverage-scored training library pack
(flare-harvest-library v1 JSON) you can hand to any training pipeline.
Or add it to Claude Desktop (claude_desktop_config.json):
With optional tokens for higher rate limits:
Requirements: Node.js β₯ 20.
You: find me training data for a small vision-language model
harvest.search
{ "topic": "vision language instruction tuning" }β ~40 hits: LLaVA-NeXT data, the Cauldron, DOCCI, LLaVA repoβ¦ plus the 12-expert atlas sweep for breadth (code, math, safetyβ¦)
harvest.compose
{ "goal": "small vision-language model", "catalog": [β¦] }β{ "title": "small vision-language model", "picks": [ { "id": "lmms-lab/LLaVA-NeXT-Data", "expert": "vision", "why": "Anchor Vision on Hugging Face" }, β¦ ], "rationale": "Local mixer staged 22 items across 12/12 experts. No mouth required.", "engine": "local" }
harvest.pack
{ "goal": "small vision-language model", "name": "VLM Mix v1", "catalog": [β¦] }βflare-harvest-libraryv1 JSON withcoverage: { filled: 12, totalExperts: 12, balance: 0.94 }and per-itemingest: { type: "hf-dataset" | "github-repo", ref }pointers.
The second half β the hands. Packs start as pointers; pull turns them into files:
harvest.resolve
{ "item": { "id": "openai/gsm8k", "source": "huggingface" } }β the loot preview: 4 parquet shards + README, each with a one-line reason, nothing stupid
harvest.pull
{ "pack": {β¦}, "destination": "download" }β resolves + fetches every item, updates the pack'sindexstanza, and buildsVLM Mix v1.harvest.zipβ one file with everything plusMANIFEST.jsonprovenance
harvest.push
{ "bundle": "/path/to/VLM Mix v1.harvest.zip", "owner": "you", "repo": "my-data", "token": "ghp_β¦", "create_if_missing": true }β one commit on a new private repo: files + manifest, message says what it is and where it came from
The same loop works for code. Second session β the build pack:
You: I want the rate-limiting code from a few good repos, packaged so my agent can build with it
harvest.search
{ "topic": "express rate limiting middleware" }β thecodeexpert surfaces the repos with the parts you want
harvest.pack
{ "goal": "add rate-limiting to my API", "name": "Rate Limit Parts v1", "catalog": [β¦], "purpose": "code" }βflare-harvest-codepackv1 JSON β same item envelope, no training-coverage scoring
harvest.pull
{ "pack": {β¦}, "destination": "download", "kinds": ["code", "docs"] }β pulls just the source files (skips weights/datasets), builds the.harvest.zip
harvest.push
{ "bundle": "β¦", "owner": "you", "repo": "rate-limit-parts", "create_if_missing": true }β your agent pulls the repo and builds from the parts
| Tool | What it does |
|---|---|
harvest.search | {topic, hf_token?, gh_token?} β topic search + full 12-expert atlas sweep, deduplicated |
harvest.expert | {expert, hf_token?, gh_token?} β run one expert's curated queries (12 ids: code, math, science, language, vision, audio, medical, law, knowledge, safety, affect, systems) |
harvest.latest | {} β trending: recently-updated HF datasets, hot/recent GitHub repos |
harvest.compose | {goal, catalog} β local mixer composes a balanced set; picks + rationale (the mixer itself needs no LLM) |
harvest.pack | {goal, catalog, name, purpose?: "training" | "code"} β compose β pack JSON. training (default) scores the mix and emits flare-harvest-library v1; code emits a flare-harvest-codepack v1: a parts bin an agent pulls and builds from |
harvest.resolve | {item, kinds?, hf_token?, gh_token?} β pull-list preview for one pack item: which files, why, and needs_token for gated items. kinds (e.g. ["code","docs"]) previews only those file kinds. No downloading |
harvest.pull | {pack, destination: "local" | "download", kinds?, dir?, confirm_large?, hf_token?, gh_token?} β resolve + fetch + index update for the whole pack; "download" also builds a .harvest.zip. kinds (e.g. ["code","docs"]) is the code-pack flow: just the source files, no weights/datasets. Oversized pulls return needs_confirm first |
harvest.push | {bundle, pack_name?, owner, repo, token, create_if_missing?, branch?, path?} β push a bundle to GitHub in one commit (private by default when creating) |
catalog items are { id, source, expert, description? } β the objects harvest.search returns drop straight in.
Everything hits the public Hugging Face and GitHub APIs. Unauthenticated, you get:
/api/datasets (occasional 429s on bursts; results are cached 12 min in-process)For real use, set HF_TOKEN / GITHUB_TOKEN (free accounts) via env or the per-call
hf_token / gh_token args. See .env.example.
Packs are flare-harvest-library v1 JSON β see PACK-FORMAT.md.
A fresh pack is pointers, not data: each item names its source and how to
ingest it, but no files. harvest.pull turns pointers into files and records
the result in an additive index stanza on the pack JSON:
v1 readers ignore the stanza safely; packs without it load as pointer-only.
destination: "download" additionally produces a .harvest.zip containing the
files plus MANIFEST.json (source, URL, license, SHA-256 per file β licenses
are reported as unknown when the pack doesn't state one, never guessed).
Tokens are passed per call, never stored:
harvest.search / harvest.expert / harvest.latest / harvest.resolve / harvest.pull
accept hf_token / gh_token (or the HF_TOKEN / GITHUB_TOKEN env vars) β
used for higher rate limits and gated datasets, sent only as an
Authorization header.harvest.push takes a per-call token (GitHub PAT, or GITHUB_TOKEN env)..harvest-fetch.json) hold only paths, byte counts,
and hashes.src/harvest/scan-core.ts β parallel HF/GitHub scanning, 12-min result cachesrc/experts.ts β the 12-expert atlas (curated queries + seed items)src/harvest/compose-local.ts β the local mixer ("No mouth required"): anchors one item per expert, then balances HF/GitHub sources, then fills for coveragesrc/coverage.ts β entropy-based coverage/balance scoring across the 12 expertssrc/harvest/compose-mouth.ts β opt-in: compose via any OpenAI-compatible /v1 endpoint (LM Studio, etc.) instead of the local mixersrc/harvest/resolve.ts β the "which files" intelligence: HF Hub / GitHub tree listings classified into data, weights, config, code, docs (pure functions + thin API clients)src/harvest/fetch.ts β the hands: streaming downloads with resume, size guards, progress, SHA-256 per file; token arrives as a function arg, never storedsrc/harvest/bundle.ts β zips a fetched dir + MANIFEST.json provenance into a .harvest.zipsrc/harvest/push.ts β pushes a bundle to GitHub via the Git Data API: one commit, private by default (// TODO(GitLab) seam sketched)src/harvest/index.ts β the agent side: additive index stanza + resolvePointer / getItemFiles / queryIndexmcp/server.ts β the MCP server (stdio, low-level SDK API, hand-written schemas)Apache-2.0 β see LICENSE.
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/harvest)<a href="https://allmcps.com/mcp/harvest"><img src="https://allmcps.com/api/badge/harvest?style=directory" alt="Harvest on AllMCPs" /></a>