# inferbench [Health: Active]

**Category:** 💻 Developer Tools  
**Repository:** https://github.com/RudrenduPaul/InferBench  
**GitHub Stars:** 0  
**npm Downloads (last month):** 213  
**Views:** 0  
**Installs:** 0  
**Upvotes:** 0  
**Directory Page:** https://allmcps.com/mcp/inferbench

## Description
Benchmarks local LLM inference speed (tokens/sec) on your own hardware via MCP tools.

## Claude Desktop Quick Installation
Install path detected from listing signals. Uses `npx` (confidence: high):

```json
"mcpServers": {
  "inferbench": {
    "command": "npx",
    "args": ["-y","inferbench-cli"]
  }
}
```

## Documentation

## What inferbench MCP server does

The inferbench MCP server is intended for measuring local language-model inference on a developer’s own hardware. It supports two engines: llama.cpp and omlx. Results report average, minimum, and maximum tokens per second across eight measured prompts, along with the number of samples used.

The benchmark is useful when choosing between supported local engines, checking how a model performs on a particular machine, or collecting repeatable throughput data for automation. Its recommendation is limited to the engines and model tested in that run; it is not a general ranking of inference systems.

## How it works

Before measurement, InferBench sends one warm-up completion to absorb first-request overhead. It then sends a fixed set of eight varied prompts to each selected engine and times the complete response body rather than stopping when response headers arrive. This produces a comparable measurement path for both engines.

Each engine is started through its own OpenAI-compatible HTTP service: llama.cpp uses `llama-server`, while omlx uses `omlx serve`. The same timing approach and prompt sweep are used for each selected engine. A run can target one engine or all installed supported engines.

The inferbench MCP server can expose the benchmarking operation to an agent, while the underlying project also provides command-line packages for JavaScript/TypeScript and Python. Reports can be emitted as human-readable tables or as camelCase JSON for scripts and CI workflows.

## Setup and configuration

The project publishes an npm package and a PyPI package named `inferbench-cli`. The npm distribution requires Node.js 18 or newer; the Python distribution requires Python 3.10 or newer. At least one supported inference engine must already be installed because InferBench does not install engines itself.

For llama.cpp, the model argument can identify a Hugging Face repository and quantization, allowing llama.cpp to download and cache the model. For omlx, the model must already exist under its local model directory; the model argument identifies that directory name. Consequently, testing both engines with one model requires that model to be available in each engine’s expected format.

The CLI accepts a required model, an optional comma-separated engine list, a maximum completion-token limit, JSON output, an output file, and verbose engine logs. Relative output paths that resolve outside the current working directory are rejected.

## Tools and capabilities

The inferbench MCP server’s supported purpose is local inference benchmarking. The project’s documented command-line and Python surfaces provide these related capabilities:

- Run benchmarks against llama.cpp, omlx, or both.
- Measure average, minimum, and maximum tokens per second.
- Save machine-readable benchmark reports.
- Detect hardware details through the Python library.
- Compare recognized models with a dated static cloud-price reference through the Python library.

The cloud comparison is not a live price lookup and returns no value for an unrecognized model.

## Limitations and notes

Only llama.cpp and omlx are documented as supported engines. Omlx is documented for Apple Silicon, and its model workflow differs from llama.cpp’s download behavior. Results depend on the selected model and the local hardware, so they should not be treated as universal engine rankings.

The supplied material documents CLI and Python interfaces but does not specify an MCP transport, MCP configuration file, environment variables, or named MCP tool identifiers. The inferbench MCP server should therefore be configured only according to additional project or host documentation that supplies those details.

_Full upstream README: https://allmcps.com/mcp/inferbench/readme_

