The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the MCP Turboquant listing page.
Self-contained Python MCP server for LLM quantization. Compress any HuggingFace model to GGUF, GPTQ, or AWQ format in a single tool call.
No external CLI required -- all quantization logic is embedded.
Or run directly with uvx:
The info, check, and recommend tools work out of the box. For actual quantization, install the backend you need:
Add to ~/.claude/settings.json:
Or with uvx (no install needed):
Add to claude_desktop_config.json:
| Tool | Description | Heavy deps? |
|---|---|---|
info | Get model info from HuggingFace (params, size, architecture) | No |
check | Check available quantization backends on the system | No |
recommend | Hardware-aware recommendation for best format + bits | No |
quantize | Quantize a model to GGUF/GPTQ/AWQ | Yes |
evaluate | Run perplexity evaluation on a quantized model | Yes |
push | Push quantized model to HuggingFace Hub | No |
Once configured, ask Claude:
"Get info on meta-llama/Llama-3.1-8B-Instruct"
"What quantization format should I use for Mistral-7B on my machine?"
"Quantize meta-llama/Llama-3.1-8B to 4-bit GGUF"
"Check which quantization backends I have installed"
"Evaluate the perplexity of my quantized model at /path/to/model.gguf"
"Push my quantized model to myuser/model-GGUF on HuggingFace"
All quantization logic runs in-process. No external CLI tools needed.
MIT