The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Transcriptor MCP listing page.
Connect one server. Then ask Claude, ChatGPT or etc about a video: the transcript, the chapters, the metadata, or a single frame. It works with 11 platforms, not only YouTube.
Connect · What to ask · Widgets · Platforms · Self-host
The hosted endpoint is:
Then run /mcp and approve the sign-in in the browser. After this, claude mcp list shows ✔ Connected.
| Client | What to do |
|---|---|
| Claude (web and desktop) | Open Settings → Customize → Connectors. Select Add → Add custom connector, paste https://transcriptor.gateway.mcpal.io/mcp, then select Add. |
| ChatGPT | Open Settings → Security and login and turn on Developer mode. Then open Plugins, select +, and paste https://transcriptor.gateway.mcpal.io/mcp. |
| Codex | Add the block below to ~/.codex/config.toml, then run codex mcp login transcriptor. The CLI, the IDE extension and the ChatGPT desktop app share this file. |
Note: ChatGPT developer mode is available on the web, for paid plans. Some releases show this control as Settings → Apps & Connectors → Advanced.
Note:
codex mcp addregisters stdio servers only, so a hosted server goes intoconfig.toml. See the Codex MCP docs.
If your client is not in the list above, add the server with this configuration:
If you want to run the server yourself, read Self-host. The tools are the same and you need no account.
| Ask for this | Tool |
|---|---|
| "Summarize this video for me" | get_transcript |
| "Give me the subtitles as an SRT file" | get_raw_subtitles |
| "Is there a German track for this video?" | get_available_subtitles |
| "Who published this and how many views?" | get_video_info |
| "Go to the part about pricing" | get_video_chapters |
| "Show me the screen at 4:12" | get_video_frame |
| "Get transcripts for the first 5 videos in this playlist" | get_playlist_transcripts |
| "Find recent videos about X" | search_videos (YouTube) |
Long transcripts come in parts. Each response gives a cursor for the next part, so no text is lost.
Each tool that takes a video accepts url. This is a link from a supported platform or a plain YouTube ID. Each tool returns content (text for the chat) and structuredContent (typed JSON for your code).
get_transcriptClean plain text, without timestamps, HTML, or speaker names. The tool finds the type and the language for you.
Response: videoId, type, lang, text, is_truncated, total_length, start_offset, end_offset. When more text is available, the response also has next_cursor.
get_raw_subtitlesRaw SRT or VTT content, in parts.
Input:
type — official or autolang — a language coderesponse_limit — default 50000, minimum 1000, maximum 200000next_cursor — the cursor of the previous responseResponse: the fields of get_transcript, plus format (srt or vtt) and content.
get_available_subtitlesResponse: official and auto. Each field is a sorted list of language codes. Use this tool first, then give type and lang to the tools above.
get_video_infoExtended metadata from yt-dlp:
videoId, title, description, webpageUrluploader, uploaderId, channel, channelId, channelUrlduration, uploadDate, viewCount, likeCount, commentCounttags, categories, liveStatus, isLive, wasLive, availabilitythumbnail and thumbnailsget_video_chaptersResponse: chapters. Each item has startTime, endTime, and title. When the video has no chapters, the list is empty.
get_video_frameInput:
timecode — "MM:SS" or "HH:MM:SS.mmm"seconds — an alternative to timecode. Give one of the two, not bothformat — jpeg (default) or pngwidth — default 1280, maximum 1920, never larger than the sourcequality — 2 to 31, for jpeg onlyResponse: an image block, plus timestampSeconds, timestamp, mimeType, sizeBytes, and width. This tool needs ffmpeg. The Docker image includes it.
get_playlist_transcriptsInput:
url — a playlist URL, or a watch URL with list=type, lang, format — the same as get_raw_subtitlesplaylistItems — a yt-dlp -I value such as 1:5, 1,3,7, or -1maxItems — the maximum number of videosResponse: results. Each item has videoId and text.
search_videosInput:
query — the search textlimit — default 10, maximum 50offset — the number of results to skipuploadDateFilter — hour, today, week, month, or yearresponse_format — json (default) or markdownResponse: results. Each item has videoId, title, url, duration, uploader, viewCount, and thumbnail.
Four tools have an interactive interface: get_transcript, get_video_info, get_video_frame, and search_videos. Clients that support MCP Apps and the ChatGPT Apps SDK show this interface in the chat. Other clients get the same data as text and JSON.
|
|
|
|
YouTube · Twitter/X · Instagram · TikTok · Twitch · Vimeo · Facebook · Bilibili · VK · Dailymotion · Reddit
Each tool that takes a video accepts a link from these 11 platforms. The tool search_videos works with YouTube only, through yt-dlp ytsearch.
The server does not download video or audio files for you. It returns text, metadata, and single frames.
The tools are the same as on the hosted endpoint. You need no account.
Run the server with Docker. The image serves Streamable HTTP on port 4200:
Then point your client at http://localhost:4200/mcp.
For stdio, give the image an explicit command:
The server starts with no environment variables. Each variable below is optional.
| Variable | Default | Function |
|---|---|---|
MCP_PORT and MCP_HOST | 4200 and 0.0.0.0 | The HTTP listener |
COOKIES_FILE_PATH | — | A Netscape cookies file for videos that need an account. See cookies.example.txt |
WHISPER_MODE | off | Set local or api to transcribe the audio when a video has no subtitles. Then set WHISPER_BASE_URL or WHISPER_API_KEY |
CACHE_MODE | off | Set redis and CACHE_REDIS_URL to cache subtitles and metadata |
YT_DLP_* | — | Timeouts, proxy, and JS runtimes. See .env.example |
The same port serves GET /health and GET /metrics. The metrics are in Prometheus format and include the mcp_* counters.
Transport. The server accepts POST /mcp only. GET and DELETE return 405. The server is stateless and sends no Mcp-Session-Id.
The Node process does not check bearer tokens. Put a reverse proxy or a gateway in front of it for authentication and TLS. The hosted endpoint works this way.
REST API. A second image gives the same extraction over plain HTTP:
The Swagger interface is at http://localhost:3000/docs. For a full stack with the API and the MCP server, read docker-compose.example.yml.
Development.
You need Node.js 20 or later, and yt-dlp in your PATH. Frame capture also needs ffmpeg. Other scripts: lint, type-check, format, test:coverage, test:e2e:api, and test:e2e:mcp.
Releases. The version comes from package.json at runtime, through src/version.ts. Change this version, move the [Unreleased] entries of the changelog into the new version, then push a v* tag. CI builds both images and publishes the MCP Registry entry from server.json.
Layout. src/mcp.ts (stdio entry), src/mcp-http.ts (Streamable HTTP), src/mcp-core.ts (tools, prompts, widgets), src/youtube.ts (yt-dlp), src/whisper.ts, src/cache.ts, src/index.ts (REST API), load/ (k6), and src/e2e/ (Docker smoke tests).
Pull requests are welcome. Fork the repository, make a branch, and make sure that npm test and npm run lint pass. Then open a pull request.
The hosted endpoint at transcriptor.gateway.mcpal.io is governed by the Terms of Service and the Privacy Policy.
A server you host yourself is not covered by those documents. It is governed by the MIT License only.
MIT © 2026 samson-art. Read LICENSE.