The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Voicely listing page.
Free, open-source, offline dictation and transcription for macOS — with a local MCP server that gives Claude Code, Codex and Cursor ears.

Option+Space, speak, press it again — the text lands wherever your cursor is, terminals and VS Code included.voicely connect registers a local MCP server, so your agent can transcribe files and read your calls and dictations.macOS 14+ on Apple Silicon. Like Superwhisper, MacWhisper or Wispr Flow — but free and MIT-licensed, 100% on-device, and built to hand text to your AI agent.
Tools that do this usually ask for a monthly subscription or a one-time payment. Voicely is free, forever: MIT-licensed, and every line of it is in this repository. It is made by one person, in one room, who uses it every day. If it saves you time, a ⭐ helps others find it.
One command. There is no DMG to download and nothing to drag into Applications:
Or through the Homebrew tap: brew install --cask stulkovld/voicely/voicely.
The script downloads the current build, verifies its SHA-256 and its signing certificate in a private staging directory, installs it transactionally — the previous version stays as a backup until the new one passes every gate, and any failure rolls back cleanly — then launches the app. Re-running it updates in place and keeps your transcripts, model and settings.
The public build is signed with Voicely's own certificate rather than an Apple Developer ID, and it is not notarized by Apple. The certificate is the same for every release, so macOS keeps your Microphone and Accessibility permissions when you update. If an older copy left permission entries behind, the installer clears them and Voicely asks again by itself.
First launch
Requirements: macOS 14 or newer on Apple Silicon.
Voicely moves what you said to where you work. It never rewrites, summarizes or "improves" what you meant — word-level truth is kept next to every transcript.
Option+SpacePress the hotkey, speak, press it again. A floating glass pill shows the live waveform while you talk. The text lands wherever your caret is: native Mac fields through Accessibility, checked by reading the caret back; Chrome, VS Code, Claude and every other Chromium or Electron app through the clipboard and ⌘V — your clipboard is put back right after the app has read the text, and the transcript is marked so clipboard managers don't keep it; terminals through typed text. If no app takes the text, it waits on the clipboard and the pill says "Press ⌘V". The Output menu can pin dictation to the clipboard permanently; secure fields are save-only, always. Dictations are saved to ~/Documents/Voicely/dictations/; where each one went, and how, is logged without the words in ~/Library/Logs/Voicely/insertion.jsonl.
Calls (menu bar → Record Call). Voicely records system audio (through ScreenCaptureKit — no virtual audio device needed) and your microphone, mixes them into one track, transcribes the mix and diarizes it globally. An echo can't duplicate a sentence, and several people on one mic separate into distinct voices. "You" is identified by matching diarized voices against your microphone's own activity. The transcript reads as speaker turns and is saved to ~/Documents/Voicely/calls/<id>/ as markdown, word-level JSONL and WAV.
Files. Any audio or video file becomes a transcript, with optional speaker diarization: voicely transcribe <file> from the terminal, or the transcribe_file tool from an agent. Results land in ~/Documents/Voicely/files/.
Either way the result is text in one tap. Hand it to your agent yourself — paste it into Claude Code, save it as a note — or let the agent read it through MCP, below.
The voicely CLI and a local MCP server make every capability agent-native. After installing the app, expose the CLI on your PATH and register the server with the agents you use:
Claude Desktop and other MCPB-aware clients can install the server in one click from voicely-<version>.mcpb on the latest release. Voicely is listed in the official MCP Registry as io.github.StulkovLD/voicely.
Teach the agent the whole playbook with the skill:
voicely mcp speaks standard stdio MCP (JSON-RPC 2.0, fully offline) and exposes four tools:
| Tool | What it does |
|---|---|
transcribe_file | Transcribe an audio/video file (optional speaker diarization) |
list_transcripts | List saved transcripts (dictations / calls / files), newest first |
get_transcript | Read a transcript by id or alias (last, last-call, …) |
get_last_call | Read the most recent call transcript |
voicely connect uses each harness's own mcp add command and never hand-edits a config file it doesn't own. Restart the agent afterward and ask it: "transcribe ~/Downloads/interview.m4a and pull the action items", "look at my last call and draft a project plan", "what did I dictate earlier?". The skill covers bootstrap from zero, online video via yt-dlp (YouTube / TikTok / Instagram) and "watching" a video by pairing word-level timestamps with extracted frames. Everything stays on the machine.
Manual config — register voicely mcp as a stdio server. The exact file differs per harness; the shape is always command: voicely, args: ["mcp"]:
| Harness | Where | How |
|---|---|---|
| Claude Code | plugin or claude mcp add | claude mcp add voicely -- voicely mcp (or the bundled plugin) |
| Codex | ~/.codex/config.toml | [mcp_servers.voicely]command = "voicely"args = ["mcp"] |
| Cursor | user settings | cursor --add-mcp '{"name":"voicely","command":"voicely","args":["mcp"]}' |
| Hermes | ~/.hermes/config.yaml | hermes mcp add voicely --command voicely --args mcp |
| OpenClaw | openclaw.json | openclaw mcp set voicely --command voicely --args mcp |
| Anything else | its MCP config | {"mcpServers":{"voicely":{"command":"voicely","args":["mcp"]}}} |
There is no daemon: the server loads the speech model itself on the first transcribe_file call.
macOS 14+ on Apple Silicon and the Xcode Command Line Tools (xcode-select --install):
In SPM the CLI product is VoicelyCLI, not
voicely: on case-insensitive APFS avoicelyproduct would collide with the app binaryVoicely. Thevoicelycommand is created bysetup, in a directory with no such collision.
Parakeet TDT 0.6B v3 (NVIDIA, CC-BY-4.0), running on-device through FluidAudio's CoreML port: about 470 MB on disk, 25 European languages including English and Russian — switching languages mid-sentence works — natively punctuated, roughly 30× realtime on Apple Silicon. Speaker diarization uses FluidAudio's pyannote community-1 segmentation and WeSpeaker embeddings, also on-device. Models download on first use; no Hugging Face account or token is required.
Everything is plain files under ~/Documents/Voicely/ — readable by you, your agent and any other tool.
Dictation (dictations/):
Call (calls/<id>/transcript.md): one block per speaker turn; your own microphone is always You. Word-level timing lives in transcript.jsonl next to it; the audio is kept as mix.wav (what was transcribed) plus the raw mic.wav and system.wav.
Stack: Swift 6 + AppKit + FluidAudio (Parakeet TDT v3 CoreML + diarization) + ScreenCaptureKit (system audio). All processing happens locally. No data leaves your machine.
System Settings → General → Login Items → + → add Voicely.app.
MIT