Search YouTube transcripts and read a video's frames; answers cite clickable timestamps.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
We haven't yet run this listing's install command through our automated sandbox check. This isn't a red flag β we're steadily working through the catalog.
π‘ Paste the JSON block into your client's configuration file under mcpServers, then restart the application.
A hosted Model Context Protocol server that lets an AI agent read YouTube videos β and cite the exact second it got the answer from.
A language model cannot watch a video. Point it at this endpoint and it gains nine tools for searching transcripts, reading a video's frames β slides, charts, demos, on-screen text β and answering questions with citations that are verified before you see them.
No integration code. No scraping. No proxy pool.
Remote-only and hosted β there is nothing to install or self-host. This repository is the public manifest, configuration reference and issue tracker for that endpoint.
Most clients need no token at all. The server speaks OAuth, so the client registers itself, sends you to VidWords to sign in, and stores a credential it refreshes on its own. You can create the account during that sign-in step. The free plan includes monthly credits and 10 Watch minutes, so you can wire this up and use it before paying anything.
Add this as a custom connector:
The host registers itself, sends you to VidWords to sign in, and shows a consent screen naming exactly what it is asking for. Registration alone grants nothing β access begins only when a signed-in person clicks Approve, and live connections can be revoked from your API page with immediate effect.
Then type /mcp in a session and choose Authenticate.
.cursor/mcp.jsonCursor shows the server as Needs login β click that once and it runs the OAuth flow in your browser. Because this file carries no secret, it is safe to commit, which the header form below is not.
For CI, a container, or a client with no OAuth support, authenticate with a header. Create an account at vidwords.com/register, verify your email, then copy the token from your profile.
claude_desktop_config.json.cursor/mcp.jsonKeep this out of version control, or use ~/.cursor/mcp.json instead β the header holds a live
credential.
~/.codex/config.tomlDo not use
bearer_token_env_var. It is the obvious-looking field, but it sendsAuthorization: Bearer <value>and this server authenticates with Basic.
This repository also ships a small stdio proxy (src/index.js) that speaks MCP on
stdin/stdout and forwards tool calls to the hosted endpoint. Use it when your client cannot
send a custom HTTP header, or when you want the server in a container:
Run straight from this repository β the proxy is not published to npm, so a bare
npx @vidwords/mcpwill not resolve.
The tool schemas are declared inline in the proxy, so initialize and tools/list answer
without any credentials and the upstream is not contacted until a tool is actually called.
A call without VIDWORDS_API_TOKEN returns a readable error rather than failing the
handshake. VIDWORDS_MCP_URL overrides the endpoint if you are pointing at a non-production
instance.
The generic mcp-remote bridge works too:
Ready-made config files live in examples/.
| Tool | What it does | Cost |
|---|---|---|
search_transcript | Find where a video discusses something. Takes one video or a list of up to 25, so one call can answer a question across a whole channel. Returns the matching moments with timestamps, quoted context, and youtube.com/watch?v=β¦&t=β¦s deep links. | 1 credit per video |
get_transcript | Full transcript text for up to 25 videos in one call. | 1 credit per video |
list_channel_videos | Resolve a channel handle, URL or UC⦠id to its recent uploads. | Free · Starter and up |
list_watchlists | The account's Radar watchlists and how much each has recorded. | Free |
watchlist_activity | Newest uploads Radar has recorded for one watchlist. | Free |
account | Plan and remaining credits, so the agent can price a job before running it. | Free |
analyze_video | Start a frame-level analysis β slides, charts, demos and on-screen text, not just captions. Returns an analysisId immediately. | Watch minutes |
get_analysis | Read a finished analysis: chapters, key points, timestamped evidence. | Free |
ask_video | Ask a question against a finished analysis. Citations are verified against stored evidence or dropped. | 1 Watch question |
search_transcript over get_transcriptBoth cost one credit per video, so there is no billing reason to choose. The reason is context.
Ask "what did this two-hour interview say about pricing?" and get_transcript returns roughly
20,000 words, of which perhaps 300 are about pricing β those 300 now compete for attention with
19,700 that are not, and the answer gets worse, slower and more expensive to generate.
search_transcript returns only the matching stretches, each with a deep link. Reach for
get_transcript when you genuinely want the whole text: an export, a diff, a corpus.
Both transcript tools take optional from and to timecodes β seconds (615), m:ss
(10:20) or h:mm:ss (1:02:13):
These are the same formats the tools print back, so a timestamp out of one answer can be pasted straight into the next question. A timecode that cannot be parsed is refused before anything is fetched, so a typo costs no credit β it never silently widens to the whole video.
search_transcript accepts a list, which is how you answer "what has this channel said about
X" without a round trip per video. Get the ids from list_channel_videos first:
Each video is billed at the usual 1 credit, and one unavailable video is reported in its own row rather than failing the call β the others were fetched and charged for, so you still get them.
analyze_video looks at slides, charts, code samples and on-screen text that is never spoken
aloud. ask_video then answers against that stored analysis, and every citation is checked
before you see it: a visual claim has to match a frame that was actually recorded, a spoken
claim has to land on a real transcript segment. Anything that fails is dropped, and when nothing
survives the answer says the evidence is insufficient rather than producing a confident guess.
That is occasionally annoying β a refusal is a worse demo than a fluent answer β and it is the only version of this feature that is safe to put in front of an agent, because an agent repeats what it is told without the scepticism a human reader applies.
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/vidwords-youtube-2)<a href="https://allmcps.com/mcp/vidwords-youtube-2"><img src="https://allmcps.com/api/badge/vidwords-youtube-2?style=directory" alt="VidWords YouTube on AllMCPs" /></a>