Tabular data retrieval. Index CSV/Excel, query rows, aggregate. 99%+ savings vs raw file reads.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
We haven't yet run this listing's install command through our automated sandbox check. This isn't a red flag β we're steadily working through the catalog.
π‘ Paste the JSON block into your client's configuration file under mcpServers, then restart the application.
jDataMunch is an MCP server for coding agents and analysts that answers questions about CSV, Excel, Parquet, and JSONL files without pasting the rows into the context window.
Index a dataset once, then retrieve column profiles, filtered rows, server-side aggregations, and cross-dataset joins β so a million-row file costs thousands of tokens instead of millions.
Install Β· Quickstart Β· Benchmarks Β· Commercial licensing
Free for personal use. Commercial use requires a paid license β terms below.
The problem. The default way an agent explores a spreadsheet is to paste it into the prompt. A 255 MB CSV with a million rows costs roughly 111 million tokens that way, and the model still has to reason through a million rows to answer "what columns are in here?"
The mechanism. jDataMunch profiles the file once β columns, types, cardinality, null rates, distributions β and stores that locally. Queries then run against the data, not against a copy of it in the prompt: filters, aggregations, and joins execute server-side and return only results.
The outcome. Orientation questions are answered from the profile. Row-level questions return matching rows. The raw file never enters the context window.
Measured on a real public dataset, not estimated. Full harness and per-query results in benchmarks/.
Corpus: LAPD crime records β 1,004,894 rows, 28 columns, 255 MB Baseline: 111,028,360 tokens to paste the raw file
describe_dataset: ~3,849 tokens β a 25,333Γ reduction Methodology & harness Β· Full results
| Task | Without jDataMunch | With jDataMunch | Reduction |
|---|---|---|---|
| Understand a dataset's shape | Paste 111M tokens | describe_dataset β ~3,849 tokens | ~25,000Γ |
| Schema + one column deep-dive | Paste 111M tokens | describe_dataset + describe_column β ~4,400 tokens | ~25,000Γ |
| Filter to matching rows | Load all 1M rows | get_rows with filters β matching rows only | ~99%+ |
| Count by category | Return all rows, aggregate in the model | aggregate(group_by=[...]) β 21 rows | ~99.9% |
What these numbers are and are not. The reduction is measured against pasting the complete file, which is what a naive agent does and what the token bill reflects. It is not measured against a competent human analyst who would never paste a 255 MB CSV. The multiple scales with file size: a 200-row spreadsheet has far less to save, and the honest figure there is closer to "no meaningful difference."
Typical latencies from the same run: describe_column on a single column, 22β33 ms and ~600 tokens.
Requirements: Python 3.10+, any MCP-compatible client.
There is no install step. jdatamunch-mcp is a stdio MCP server with no CLI subcommands, so nothing needs to land on your PATH β point your client at uvx and it fetches and runs the server on demand.
Claude Code setup:
Nothing else. Don't have uv yet?
Reading Excel or Parquet? Those pull optional extras, which uvx takes on the --from argument:
| Command | Use it when |
|---|---|
uv tool install jdatamunch-mcp | You want it resolved once instead of per-launch |
pipx install jdatamunch-mcp | You already standardise on pipx |
pip install jdatamunch-mcp | Inside a virtualenv you manage yourself. β Refused on PEP 668 distros (Ubuntu 24.04+, Debian 12+) β use one of the two above. |
Extras take the usual bracket form here: uv tool install "jdatamunch-mcp[excel,parquet]". Registering the server still works the same way; substitute jdatamunch-mcp for uvx jdatamunch-mcp in the claude mcp add line above.
Restart Claude Code, then type /mcp β jdatamunch should be listed. That listing is the verification step; running the server directly just waits on stdin.
Full per-client setup, including Claude Desktop, Cursor, and Windsurf: QUICKSTART.md.
Assumes: jDataMunch installed and registered with your client, and a CSV to hand.
Everything happens inside your agent β there is no separate indexing command. Ask it to index:
Using jdatamunch, index ./data/sales.csv
It calls index_local, which returns the dataset name, row and column counts, and detected types. Then:
Using jdatamunch, describe the sales dataset and tell me which columns have missing values.
The agent calls describe_dataset, which returns column names, inferred types, cardinality, null rates, and sample values β without reading a single row into context. _meta.tokens_saved reports what that cost against loading the file.
Next step: describe_column for a distribution on one column, or aggregate to group and count server-side.
describe_dataset, describe_column, sample_rows, get_distribution, get_correlations.get_rows with filters, aggregate with group_by, run_sql, and plan_query to preview cost before running.suggest_joins, suggest_keys, join_datasets.get_dataset_health, data_health_radar, get_data_hotspots (null rate, cardinality anomalies, outlier spread), get_schema_drift, find_unused_columns.check_column_drop_safe and get_schema_impact before you drop or rename.search_data and find_similar_columns when you know what you mean but not what it is called.index_repo pulls CSV, Excel, Parquet, and JSONL straight from a repository, incrementally by HEAD SHA, private repos included.39 tools in total. Full reference: USER-MANUAL.md.
Everything runs locally. The dataset is profiled on your machine and the index is stored on your machine; no hosted service is involved in indexing or querying.
Aggregations and filters run against the stored data rather than being simulated in the model, which is why the row count barely affects the token cost of an answer. Sampling-based statistics report their error bounds (roughly 2% standard error) rather than presenting an estimate as exact.
| Format | Extensions | Install extra |
|---|---|---|
| CSV / TSV | .csv, .tsv | built in |
| JSON Lines | .jsonl | built in |
| Excel | .xlsx, .xls | pip install "jdatamunch-mcp[excel]" |
| Parquet | .parquet | pip install "jdatamunch-mcp[parquet]" |
Local-first. Your data is profiled and indexed on your machine and is not uploaded.
The base package's only default network behavior is an anonymous savings counter β a random ID plus aggregate token counts. No data, no column names, no file paths, no PII. Opt out completely:
index_repo reaches GitHub only when you invoke it, using a token you supply. Embedding providers are called only when you configure one. There is no scheduler and no background reporting.
Full detail, including what each optional extra pulls in: SECURITY.md.
describe_column will not be labelled offloadable. jDataMunch does not assert index freshness it cannot prove, so the cheap freshness reading answers unknown and the annotation fails closed. That is deliberate β see the annotation section.JMUNCH_OFFLOADABLE=1 (suite-wide) or JDATAMUNCH_OFFLOADABLE=1 (this server only) makes describe_column carry an advisory _meta.offloadable block marking whether the answer is simple and self-contained enough to hand to a cheaper model.
It is a label and nothing else. jDataMunch never calls another model, never routes the request, and never touches your API keys. Off by default; you decide what happens next.
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/jdatamunch-mcp)<a href="https://allmcps.com/mcp/jdatamunch-mcp"><img src="https://allmcps.com/api/badge/jdatamunch-mcp?style=directory" alt="JDataMunch MCP on AllMCPs" /></a>