The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Vox listing page.
A native macOS MCP server that gives Claude a voice.
Built in Swift. No Node.js. No Python. No Electron. Just a single binary.
Then just say: listen — and Claude hears you.
| Tool | What it does |
|---|---|
listen | Activates the mic, waits for you to speak, returns transcript + detected language when silence is detected |
speak | Speaks text aloud — ElevenLabs TTS with automatic fallback to macOS system voice |
ELEVENLABS_API_KEY in ~/.claude/.env for high-quality multilingual TTSCopy and paste this into Claude Code:
Claude will handle the download, configuration, and restart.
Then add to ~/.claude.json:
Restart Claude Code.
Add to ~/.claude.json:
Restart Claude Code.
Default TTS voice is Hana via ElevenLabs (eleven_multilingual_v2). Supports Chinese, English, Czech, Vietnamese, and 20+ languages in the same voice. Without an ElevenLabs key, falls back to macOS system voice automatically.
No ports. No sockets. No daemon. Just stdin/stdout.
"No speech detected" — speak within ~1s of calling listen, VAD cuts off after 0.8s silence.
No default input device — Mac mini has no built-in mic. Connect USB/Bluetooth mic or iPhone via Continuity Camera, set default in System Settings → Sound → Input.
ElevenLabs silent — add ELEVENLABS_API_KEY=sk-... to ~/.claude/.env, or leave it out to use system voice.