Fixes the names and jargon speech-to-text gets wrong before your agent acts on a dictated prompt.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent โ or use 1-click editor setup below.
One-click editor setup isnโt available for this listing yet โ we donโt have a confirmed install command, and weโd rather show nothing than point your editor at the wrong package or host. Follow the projectโs own setup instructions, linked above.
A personal lexicon for voice-to-agents.
One YAML file of the words speech-to-text gets wrong, applied everywhere your voice lands: MCP, Claude Code, the browser, macOS.
Live demo ย ยทย Quickstart ย ยทย Docs ย ยทย lexicon.ashlr.ai
![]()
That is the live demo running this repo's real matcher on your text, in your browser, with nothing installed. (Its Dictate button uses your browser's own speech recognizer, which in Chrome sends audio to Google.)
Method, full tables and every failing case are in docs/BENCHMARK.md. Reproduce with npm run bench (no setup beyond a clone), or npm run bench:audio && npm run bench:compare (needs macOS and whisper.cpp).
Against the alternatives. 330 clips of real audio through whisper.cpp small.en, three voices. Same audio, same recognizer, same 70-term lexicon in every row; the only thing that changes is how the proper nouns get fixed.
| how the words get fixed | proper nouns recovered | clean prose wrongly changed |
|---|---|---|
| nothing, raw whisper.cpp | 45.9% | n/a, nothing runs |
| exact-string substitution, the macOS Text Replacement approach | 62.0% | 0 of 72 |
| the same, plus a casing rule per term | 71.3% | 0 of 72 |
whisper.cpp's own --prompt hint list | 76.0% | n/a, nothing runs |
| Lexicon | 91.0% | 0 of 72 |
Every row uses the same curated seventy-term lexicon (bench/lexicon.yaml). The prose column is a property of the terms in the file as much as of the matcher, so a lexicon assembled some other way is a different measurement.
Exact substitution recovers the spellings someone already wrote down, and nothing else. It cannot reach Versal, Superbase, CloudFloor or pedantic, because no table written by hand contains the mistake you have not heard yet. The phonetic and fuzzy tiers exist for that gap and recover 31 of the 279 term slots on their own, which is 11 points. The remaining 9 points over the casing-aware row come from the alias tier's tolerance for how the recognizer breaks a name into words, since that row already matches case. --prompt is a complement rather than a rival: stacked with the lexicon it reaches 95.7%.
Two honest notes about that table. Raw whisper.cpp cannot wrongly change prose because nothing runs, which is the absence of the feature rather than an advantage. --prompt is not the same case: it biases the recognizer itself, so whatever it changes is already in the transcript before scoring begins, while the metric counts sentences a post-pass altered. That cell is unmeasured rather than zero. Telling which way it goes would mean diffing prompted transcripts against unprompted ones, which this harness does not do.
Before and after.
| corpus | proper nouns recovered, raw STT | after lexicon | clean prose wrongly changed |
|---|---|---|---|
| real audio, whisper.cpp base.en (330 clips) | 41.9% | 82.8% | 0 of 72 |
| real audio, whisper.cpp small.en with prompt hints | 76.0% | 95.7% | 0 of 72 |
| synthetic STT errors (402 sentences, 70 terms) | 5.1% | 96.5% | 0 of 95 |
Latency is about 0.3 ms per sentence. The real-audio rows use macOS text-to-speech read into whisper.cpp, so they are cleaner than a phone microphone.
The synthetic 5.1% is not a claim that speech-to-text gets 5% of proper nouns right in general. Every sentence in that corpus was written to contain a mis-hearing, so 5.1% is only the handful that came out right anyway. The honest "before" number is the real-audio one, 41.9%.
The last column counts ordinary prose only. Each corpus also contains sentences deliberately built to trip the matcher (a bare "llama" next to an Ollama term, sound-alikes, code spans), marked expected-hard; with those included the false-positive rate is 15.3% (19 of 124) synthetic and 16.7% (15 of 90) on audio. Both numbers, and every failing case, are in docs/BENCHMARK.md.
The package is @ashlr/lexicon; the command is lexicon.
Then open Claude Code and say a sentence with your company name in it. Done.
The install script runs lexicon setup for you (LEXICON_NO_SETUP=1 skips it); after a Homebrew or npm install, run it yourself. Every step is optional and safe to rerun, and lexicon setup --dry-run writes nothing while describing the run you would get from the same command without it: the steps it would perform, and the ones it would stop and ask about, with the answer pressing Enter gives each.
With Homebrew, always use the full tap name ashlrai/tap/lexicon. Plain brew install lexicon installs dns-lexicon, an unrelated DNS tool in homebrew-core.
lexicon setup does, in seven numbered stepsThe full walkthrough, with the real terminal output, is in docs/QUICKSTART.md.
Or skip the wizard and add one term by hand. The first argument is the canonical spelling, the rest are what STT actually produces:
Claude Code plugin, if you would rather not install a CLI at all. No Node install step, no build:
lexicon doctor checks the install. There is no telemetry and all state is local files: the CLI, hooks, MCP server, local API and extension make no request beyond loopback. The one outbound request in the codebase is lexicon voice fetching a whisper model on first use. The install script, npm and Homebrew fetch the package itself. See SECURITY.md.
Speech-to-text is about 95% accurate on ordinary English and much worse on invented names. In the benchmark above, raw whisper.cpp base.en transcribed 117 of 279 dictated proper nouns correctly. "Ashlr.AI" becomes "Ashler", "Kubernetes" becomes "Cooper Nettie's", "SaaS" becomes "sauce", "auth" becomes "off". Those are exactly the words an agent needs to get right.
Dictation apps (Wispr Flow, Superwhisper, Aqua) each keep their own dictionary and none of them share it. Agents (Claude Code /voice, ChatGPT voice, Codex, local Whisper) run their own recognizer with no user vocabulary at all. This is the portable layer in between: corrections happen after STT and before the model, wherever the text passes through.
This is not a dictation app. It sits between whatever dictation you already use and whatever agent you talk to. The research behind that call, including the kill criteria, is in docs/RESEARCH.md.
No reviews yet โ be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/lexicon-2)<a href="https://allmcps.com/mcp/lexicon-2"><img src="https://allmcps.com/api/badge/lexicon-2?style=directory" alt="Lexicon on AllMCPs" /></a>