Benchmarks local LLM throughput on your hardware across llama.cpp and omlx using fixed prompts.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent — or use 1-click editor setup below.
We haven't yet run this listing's install command through our automated sandbox check. This isn't a red flag — we're steadily working through the catalog.
💡 Paste the JSON block into your client's configuration file under mcpServers, then restart the application.
Inspect callable tools, capabilities, and parameters exposed to AI agents by Inferbench.
The inferbench MCP server is intended for measuring local language-model inference on a developer’s own hardware. It supports two engines: llama.cpp and omlx. Results report average, minimum, and maximum tokens per second across eight measured prompts, along with the number of samples used.
The benchmark is useful when choosing between supported local engines, checking how a model performs on a particular machine, or collecting repeatable throughput data for automation. Its recommendation is limited to the engines and model tested in that run; it is not a general ranking of inference systems.
Before measurement, InferBench sends one warm-up completion to absorb first-request overhead. It then sends a fixed set of eight varied prompts to each selected engine and times the complete response body rather than stopping when response headers arrive. This produces a comparable measurement path for both engines.
Each engine is started through its own OpenAI-compatible HTTP service: llama.cpp uses llama-server, while omlx uses omlx serve. The same timing approach and prompt sweep are used for each selected engine. A run can target one engine or all installed supported engines.
The inferbench MCP server can expose the benchmarking operation to an agent, while the underlying project also provides command-line packages for JavaScript/TypeScript and Python. Reports can be emitted as human-readable tables or as camelCase JSON for scripts and CI workflows.
The project publishes an npm package and a PyPI package named inferbench-cli. The npm distribution requires Node.js 18 or newer; the Python distribution requires Python 3.10 or newer. At least one supported inference engine must already be installed because InferBench does not install engines itself.
For llama.cpp, the model argument can identify a Hugging Face repository and quantization, allowing llama.cpp to download and cache the model. For omlx, the model must already exist under its local model directory; the model argument identifies that directory name. Consequently, testing both engines with one model requires that model to be available in each engine’s expected format.
The CLI accepts a required model, an optional comma-separated engine list, a maximum completion-token limit, JSON output, an output file, and verbose engine logs. Relative output paths that resolve outside the current working directory are rejected.
The inferbench MCP server’s supported purpose is local inference benchmarking. The project’s documented command-line and Python surfaces provide these related capabilities:
The cloud comparison is not a live price lookup and returns no value for an unrecognized model.
Only llama.cpp and omlx are documented as supported engines. Omlx is documented for Apple Silicon, and its model workflow differs from llama.cpp’s download behavior. Results depend on the selected model and the local hardware, so they should not be treated as universal engine rankings.
The supplied material documents CLI and Python interfaces but does not specify an MCP transport, MCP configuration file, environment variables, or named MCP tool identifiers. The inferbench MCP server should therefore be configured only according to additional project or host documentation that supplies those details.
Factual signals from GitHub, npm, and our automated checks — not a rating.
No reviews yet — be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/inferbench)<a href="https://allmcps.com/mcp/inferbench"><img src="https://allmcps.com/api/badge/inferbench?style=directory" alt="Inferbench on AllMCPs" /></a>