The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Talkies listing page.
Self-hosted speech services in one Docker image: OpenAI-compatible file transcription and text-to-speech, Talkies live ASR over WebSocket, file staging, model lifecycle controls, and an MCP endpoint for ASR workflows.
Restrict the first boot to the models you need; otherwise the entrypoint downloads every model in the bundled registry.
For CUDA-only models — Parakeet-TDT, the larger Canary models, Qwen3 TTS and
Chatterbox Turbo — use psyb0t/talkies:latest-cuda with --gpus all. The
loopback port mapping keeps the service local; see
Getting started for first boot and authentication.
| Surface | Purpose | Reference |
|---|---|---|
POST /v1/audio/transcriptions | File transcription and subtitles | HTTP API |
WS /v1/audio/transcriptions/stream | Live 16 kHz PCM ASR | Streaming |
POST /v1/audio/speech | Speech synthesis in six formats | HTTP API |
GET /v1/models | Enabled slugs and their modality | HTTP API |
GET /v1/audio/voices | Per-model voice catalog with origin tags | Models |
GET/PUT/DELETE /v1/files/* | Server-side file staging | HTTP API |
/api/ps, /unload | Model inspection and eviction | Operations |
/v1/mcp | Streamable HTTP MCP with ASR/file tools | HTTP API |
GET /healthz | Liveness probe; the only unauthenticated route | Operations |
The HTTP transcription and speech routes use the corresponding OpenAI wire shapes where those contracts overlap. Streaming ASR, files, lifecycle controls, and MCP are Talkies extensions.
wav2vec2-xlsr-53-espeak and zipa-ipa return the IPA
phones that were spoken, not words, with no language model correcting them
toward the nearest dictionary entry. Same transcription endpoint and
timestamp options as the other ASR models; see
Phoneme recognition.response_format="pcm"; other TTS formats and Kokoro are buffered.[sigh], [whispering] and [laugh] directly in the input text. Its output
carries a neural watermark by default; set TALKIES_CHATTERBOX_WATERMARK to
false to emit unmarked audio..wav into /data/custom-voices and it appears on
GET /v1/audio/voices. Qwen3 pairs it with an optional sibling .txt
transcript; Chatterbox needs only the clip, longer than five seconds.Exact slugs, executors, tag list, and registry format: Models and registries.
wav2vec2-xlsr-53-espeak and zipa-ipa use the same transcription call as
every other ASR slug; only the model changes. text comes back as a
space-separated IPA phone stream rather than words, and no language model
corrects a mispronunciation toward a real word.
Add -F "response_format=verbose_json" (or timestamp_granularities[]=word)
to get each phone as a words entry with start and end in seconds.
Tags go inline in input, in square brackets, lowercase. They are real tokens
in the model's tokenizer, so only these 19 do anything — any other bracketed
word is spoken as literal text:
Swap "voice" for the name of any .wav you dropped in /data/custom-voices
(extension stripped) to speak the same line in a cloned voice.
| Guide | Contents |
|---|---|
| Getting started | Run CPU/CUDA, persist data, authenticate, verify |
| Models and registries | Bundled slugs, image availability, custom registries |
| Architecture | Request flow, backend selection, on-disk layout |
| HTTP API | Requests, responses, files, lifecycle, MCP |
| Streaming | Live ASR protocol, streaming backends, PCM TTS |
| Configuration | Supported environment variables and limits |
| Operations and security | Exposure, model memory, data retention, logs |
| Development | Make targets, test suites, image builds |
The Talkies skill teaches agents to use the HTTP,
WebSocket, and MCP surfaces. Install it through the shared psyb0t marketplace
or let Codex discover it directly from this checkout.
Claude Code prompts for the Talkies URL and, when enabled, the bearer token; the sensitive token is stored through the client's protected configuration.
A marketplace install invokes the skill as $talkies:talkies. Codex also
discovers .agents/skills/talkies directly in this repository, where it is
invoked as $talkies without installation.
The skill and MCP bridge are published through ClawHub:
The bridge connects local stdio MCP clients to a running Talkies /v1/mcp
endpoint. Set TALKIES_URL and, when authentication is enabled,
TALKIES_AUTH_TOKEN.
TALKIES_AUTH_TOKEN enables a shared bearer token for every HTTP and WebSocket
route except /healthz. It is unset by default. Keep the port loopback-only or
put Talkies behind TLS, authentication, and rate limiting. If untrusted callers
can supply remote file_path URLs, set TALKIES_BLOCK_PRIVATE_DOWNLOADS=true.
See Operations and security for the complete posture.
make help lists every target.
Talkies is released under the WTFPL. Model weights are downloaded at runtime and have their own terms; image component notices are in THIRD_PARTY.md. Release notes are in CHANGELOG.md.