The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the MCP Video Analyzer listing page.
Turn any video — YouTube, Instagram, TikTok, Loom, X, Vimeo, direct links, local files — into transcripts, key frames, OCR text, and metadata for AI agents.
No existing video MCP combines transcripts + visual frames + metadata in one tool. This one does — across Loom, the major yt-dlp platforms (YouTube/Vimeo/TikTok/Instagram/X/Twitch/Dailymotion/Facebook), direct video URLs, and local files.
Want a full pipeline, not just a tool? social-knowledge-base is built on top of this server — it downloads whole Instagram creator accounts (reels, stories, highlights), transcribes them, and turns the result into a searchable, RAG-queryable knowledge base with AI-generated notes. Use this MCP when you want per-video analysis inside an agent; use social-knowledge-base when you want to archive and query an entire account.
npxpip install yt-dlpWithout yt-dlp or Chrome, direct URLs and local files still get frames — the bundled
ffmpeg-staticdoes the extraction, and Loom falls back to its own CDN download. Platform URLs (YouTube etc.) degrade to a clear "install yt-dlp" warning. Transcripts, metadata, and comments never require either.
There are three ways in: the /video plugin (Claude Code — slash command + MCP server auto-configured), a plain MCP server config (any MCP client), or the portable skill + CLI (Codex, Cursor, Copilot, and any agent with a shell — no MCP required).
/video plugin (recommended)This adds the /video slash command and auto-registers the MCP server — no claude mcp add needed:
Installs the video skill (Agent Skills format) into every agent detected on your machine. Agents without the MCP server configured fall back to the bundled CLI automatically — zero configuration.
Then restart Claude Code or start a new conversation.
Add to your MCP settings file:
File → Preferences → Settings → search "MCP" or edit ~/.vscode/mcp.json / %APPDATA%\Code\User\mcp.json (Windows)Settings → MCP Servers → AddThen reload the window (Ctrl+Shift+P → "Developer: Reload Window").
Add to your Claude Desktop config file:
~/Library/Application Support/Claude/claude_desktop_config.json%APPDATA%\Claude\claude_desktop_config.jsonThen restart Claude Desktop.
The same engine is exposed as a one-shot command — this is what the video skill uses on agents without MCP, and it works standalone in any terminal:
stdout is a single JSON document — metadata, transcript, ocrResults, timeline, warnings, frameCount, and frames as { time, filePath, mimeType } entries pointing at JPEG key frames copied to --out (default: the per-user cache dir — %LOCALAPPDATA% on Windows, ~/Library/Caches on macOS, $XDG_CACHE_HOME or ~/.cache on Linux — under mcp-video-analyzer/<url-hash>/; set MCP_CACHE_DIR to an absolute path to relocate it). Unlike the temp dir this used to live in, nothing reaps that location, so frames persist until you delete them — the directories are created 0700. Progress streams on stderr, so stdout can be piped straight into a JSON parser. Partial failures land in warnings with exit code 0; only hard failures exit 1.
| Flag | Description |
|---|---|
--detail <level> | brief (metadata + transcript, no frames), standard (default), detailed |
--max-frames <n> | Max key frames, 1–60 (default adapts to duration) |
--max-width <px> | Width cap for emitted frames (default 800, or MCP_FRAME_MAX_WIDTH); 0 keeps the source resolution — see Frame size |
--fields <list> | Output filter — comma-separated subset: metadata,transcript,frames,comments,chapters,ocrResults,timeline,aiSummary. Filters the emitted JSON only; use --detail brief to actually skip download/frame extraction |
--force-refresh | Bypass the cache and re-analyze |
--ocr-language <codes> | Tesseract languages (default eng+por) |
--model <name> / --language <code> | Whisper overrides for the transcription fallback |
--out <dir> | Where frame images are copied |
Run with no arguments (npx mcp-video-analyzer@latest) to start the MCP stdio server — the CLI is purely additive.
Once installed, ask your AI assistant:
(also works with an Instagram/TikTok/Loom link, a direct .mp4 URL, or a local file path). If the server is connected, it will automatically call the analyze_video tool.
Eight tools — the AI picks the cheapest one for the job and calls it automatically. Click any tool to expand its parameters and examples.
| Tool | What it does |
|---|---|
analyze_video | Full analysis: transcript + key frames + OCR + timeline + metadata |
analyze_videos | Batch version, one structured result per source (resumable) |
get_transcript | Transcript only (native captions or Whisper fallback) |
get_metadata | Metadata + comments + chapters, no download |
get_frames | Key frames only (scene-change or dense 1 fps) |
analyze_moment | Deep-dive on a time range (burst frames + transcript + OCR) |
get_frame_at | Single frame at a timestamp |
get_frame_burst | N frames across a narrow window (motion/animation) |
analyze_video — full video analysisExtracts everything from a video URL in one call:
Returns:
The AI will automatically call this tool when it sees a video URL — no need to ask.
Options:
detail — analysis depth: "brief" (metadata + truncated transcript, no frames), "standard" (default), "detailed" (dense sampling, more frames)fields — array of specific fields to return, e.g. ["metadata", "transcript"]. Available: metadata, transcript, frames, comments, chapters, ocrResults, timeline, aiSummarymaxFrames (1-60) — cap on extracted frames. Default scales with video duration at standard detail (~12 for ≤30s up to 60 for >10min); fixed 60 at detailed, 0 at brief. An explicit value always winsthreshold (0.0-1.0, default 0.1) — scene-change sensitivityforceRefresh — bypass cache and re-analyzeskipFrames — skip frame extraction for transcript-only analysismodel / language / initialPrompt — per-call Whisper overrides for the transcription fallback (override WHISPER_MODEL / WHISPER_LANGUAGE / WHISPER_PROMPT for this call only — pick a heavier model or a domain glossary for one hard clip without restarting the server)analyze_videos — batch analysisRuns analyze_video over a list of sources with a concurrency limit (default 2), returning one structured result per source — counts + warnings on success, or a per-item error on failure (one bad file never aborts the batch). Frame images are not inlined and full transcript/OCR/timeline are returned only when fields is set; otherwise you get counts. Pair with MCP_WRITE_SIDECARS=1 (below) so each video's result persists to disk and a re-run resumes instead of recomputing.
get_transcript — transcript onlyQuick transcript extraction. Falls back to Whisper transcription when no native transcript is available. Accepts the same per-call model / language / initialPrompt overrides as analyze_video.
get_metadata — metadata onlyReturns metadata, comments, chapters, and AI summary without downloading the video.
get_frames — frames onlyTwo modes:
dense: true) — 1 frame/sec for full coverageanalyze_moment — deep-dive on a time rangeCombines burst frame extraction + filtered transcript + OCR + annotated timeline for a focused segment. Use when you need to understand exactly what happens at a specific moment.
get_frame_at — single frame at a timestampThe AI reads the transcript, spots a critical moment, and requests the exact frame to see what's on screen.
get_frame_burst — N frames in a time rangeFor motion, vibration, animations, or fast scrolling — burst mode captures N frames in a narrow window so the AI can see frame-by-frame changes.
| Level | Frames | Transcript | OCR | Timeline | Use case |
|---|---|---|---|---|---|
brief | None | First 10 entries | No | No | Quick check — what's this video about? |
standard | Duration-adaptive: ~12 (≤30s) up to 60 (>10min), scene-change | Full | Yes | Yes | Default — full analysis |
detailed | Up to 60 (1fps dense) | Full | Yes | Yes | Deep analysis — every second captured |
Results are cached in memory for 10 minutes. Subsequent calls with the same URL and options return instantly. Use forceRefresh: true to bypass the cache. skipFrames is part of the cache and sidecar key, so a transcript-only analysis and a framed one of the same URL never answer for each other.
The in-memory cache is lost on restart, which makes reprocessing a large local corpus costly. Set MCP_WRITE_SIDECARS=1 to also persist results next to each local video so the work survives restarts and can resume:
<stem>.vtt — the transcript, only when it was generated by the Whisper fallback (an existing <stem>.vtt from your own pipeline is never overwritten). A later call reuses it via the normal sidecar reader and skips Whisper entirely.<stem>.analysis.json + <stem>.frames/ — the full result (frames + OCR + timeline), keyed by the video's mtime:size and the analysis params. On a later call with a matching stamp + params, the result is returned straight from disk (no extraction, no OCR).This makes analyze_videos over thousands of files resumable, and lets an external GPU transcription pipeline and this MCP share results through the filesystem: the pipeline writes <stem>.vtt, and the MCP picks it up instead of running Whisper.
| Source | Transcript | Metadata | Comments | Frames | Auth |
|---|---|---|---|---|---|
| Loom | Yes | Yes | Yes | Yes (usually needs yt-dlp — see note) | None |
| YouTube / Vimeo / TikTok / Instagram / X / Twitch / Dailymotion / Facebook | Native captions (uploaded > auto-generated) or Whisper fallback | Yes (title, duration, uploader, views, chapters, upload date) | No | Yes (capped at 1080p) | yt-dlp installed; cookies for Instagram / age-restricted (see below) |
| Direct URL (.mp4, .mov, .mkv, .webm, …) | No | Duration only | No | Yes | None |
| Direct URL + TwelveLabs | Yes (Pegasus, best-effort) | Duration floor + title | No | Yes | TWELVELABS_API_KEY |
Local file (absolute path or file:// URI) | Sidecar .vtt/.srt or Whisper fallback | Probed via ffmpeg (duration, dims, codec, audio presence) | No | Yes | None |
Loom frames: transcript, metadata, and comments come straight from Loom's API with no extra tooling. Frame extraction is different — Loom serves most videos as separate DASH video+audio streams, which only yt-dlp (
pip install yt-dlp) fetches and merges. Merging uses the bundledffmpeg-static, so no system ffmpeg is required. Without yt-dlp a direct-CDN fallback still covers some videos; when it can't, you get transcript + metadata + comments plus a warning explaining why frames are missing.Local files: pass an absolute path (e.g.,
/Users/you/clip.mp4) or afile://URI as theurlargument to any tool. Relative paths are rejected — the server's working directory is unpredictable from the MCP client. Note that any caller of the MCP server can ask it to read any file the server process has access to. UNC / network share paths (\\host\share\clip.mp4) are the exception: they reach the network rather than local disk, so they follow the network destination rules and needMCP_ALLOW_PRIVATE_URLS=1.Sidecar transcripts: if a
clip.vtt,clip.srt,clip.en.vtt, etc. lives next toclip.mp4, it's used as the transcript automatically — no Whisper roundtrip needed. SRT is converted to VTT in-memory.Embedded subtitles: if no sidecar is found and the container has an embedded subtitle stream (common in
.mkv/.mov/.mp4from screen recorders), it's transmuxed to VTT via ffmpeg and used as the transcript.Recognized extensions (local files and direct URLs):
.mp4.mov.mkv.webm.avi.m4v.wmv.flv.mpeg.mpg.m2ts.mts.3gp.ogv. The extension only gates routing — ffmpeg does the actual demuxing, so most common containers work..tsis excluded to avoid colliding with TypeScript source files.
Only http:// and https:// URLs are fetched, and only to public addresses. Requests to loopback (localhost, 127.0.0.1, ::1), private/LAN ranges, link-local, CGNAT, .local mDNS names, and Windows UNC paths are refused, as are non-HTTP schemes like ftp:// and data:.
This matters because the url argument is attacker-reachable in the normal case: an agent driving this server can be steered by the content it reads, so a URL it passes in is not necessarily one the user chose. Without the restriction the server is a proxy into whatever network it happens to sit on.
The check runs on the resolved address, not just the text, so a public hostname that resolves to 10.0.0.5 is refused too — and every hop of a redirect chain is re-checked, since a public URL answering 302 Location: http://127.0.0.1/ would otherwise walk straight past a first-hop-only check.
Cloud instance metadata endpoints (169.254.169.254, Azure's 168.63.129.16, and friends) stay blocked even with that set — there is no legitimate video there, and they are what an SSRF is usually after.
Known limitation: the address is checked at resolution time, not at connection time, so DNS rebinding — a domain answering a public address to the check and a private one to the connection — is not covered. Run the server behind an egress proxy if that is in your threat model.
Single-video pages on major platforms route through yt-dlp (pip install yt-dlp — required for these URLs). Playlists, channels, and profile pages are rejected by design; pass individual video URLs (batch them with analyze_videos).
WHISPER_LANGUAGE (e.g. pt) is also used to pick the caption language. Videos with no captions at all fall through to the normal Whisper chain.ffmpeg-static (no system ffmpeg required).| Env var | What it does | Example |
|---|---|---|
YTDLP_COOKIES | Cookie file (Netscape format), wins when both are set | C:/secrets/cookies.txt |
YTDLP_COOKIES_FROM_BROWSER | Extract cookies from an installed browser | chrome, edge, firefox |
Browser cookie extraction requires the browser to be closed on Windows (the cookie database is locked while it runs). If that's inconvenient, export a
cookies.txtonce (e.g. with a "Get cookies.txt" browser extension) and pointYTDLP_COOKIESat it. Private/age-restricted videos without valid cookies don't crash the tool — the yt-dlpERROR:line surfaces inwarnings[].
Set the TWELVELABS_API_KEY environment variable to analyze direct video URLs with TwelveLabs Pegasus. Pegasus analyzes the video server-side (visuals and its own audio) and returns an AI-generated, timestamped transcript plus an AI summary as text — capabilities the DirectAdapter can't provide (a raw .mp4 URL has no transcript or summary on its own), and with no Whisper key required.
The transcript is best-effort LLM output, not a deterministic ASR dump: Pegasus is prompted to emit [MM:SS] line rows, and lines that don't match that shape are dropped, so wording and exact timestamps depend on the model's prompt adherence. Failures (bad key, timeout, API error) surface in the tool's warnings[] rather than silently returning an empty transcript.
The biggest win is on the text-only paths: get_transcript and get_metadata return a Pegasus transcript and summary for direct URLs — a few KB of text, no frame images, no per-frame token cost. analyze_video at detail: "standard"/"detailed" still extracts frames in addition (use detail: "brief" to stay text-only).
Long videos: the summary and full transcript share a single capped completion (
max_tokens= 16384), so for very long videos the transcript may be truncated. For multi-hour content, chunking by time window is the better approach.
It's fully opt-in and non-breaking: when TWELVELABS_API_KEY is set the TwelveLabsAdapter handles direct video URLs (it registers the public URL with TwelveLabs — no upload); when it's unset, the DirectAdapter handles them exactly as before. Loom URLs are unaffected. Get a key at playground.twelvelabs.io.
When a source has no native transcript (no sidecar .vtt/.srt, no embedded subtitles, no platform captions), the audio track is transcribed with Whisper via a graceful fallback chain (in execution order):
Silent tracks: before any Whisper run, the audio is probed with ffmpeg
volumedetect(first 2 minutes). A present-but-mute track — common in muted Reels/Stories — skips transcription entirely and emits a warning that the empty transcript is expected content, not an error, saving a pointless Whisper run.
WHISPER_HF_MODEL is explicitly set. When it's unset (the default) the strategy is skipped entirely, so the CLI below wins and its WHISPER_MODEL/WHISPER_LANGUAGE settings are never silently overridden.whisper CLI — used when a whisper executable is found (pip install -U openai-whisper). Point WHISPER_BIN at the executable if it isn't on PATH. Model via WHISPER_MODEL, language via WHISPER_LANGUAGE. The bundled ffmpeg-static is put on the CLI's PATH automatically, so no system ffmpeg is required.OPENAI_API_KEY is set.No backend configured? If none of the three is available (no
whisperonPATH/WHISPER_BIN, noOPENAI_API_KEY, noWHISPER_HF_MODEL), transcription tools return an empty transcript with a warning telling you how to enable one — rather than a silent "no transcript". Installopenai-whisperor set one of the keys above. (The CLI is spawned withPYTHONUTF8=1so non-English/CJK transcripts don't crash the Python process on Windows.)
| Env var | Applies to | Default | Example |
|---|---|---|---|
WHISPER_MODEL | whisper CLI | tiny | small, medium |
WHISPER_LANGUAGE | whisper CLI / OpenAI API | auto-detect | pt, en, es |
WHISPER_PROMPT | whisper CLI / OpenAI API | — | Doha, Smiles, Livelo, Latam, milheiro |
WHISPER_BIN | whisper CLI | whisper (on PATH) | C:/.../Scripts/whisper.exe |
WHISPER_DEVICE | whisper CLI (sent only if set) | — | cuda, cpu |
WHISPER_COMPUTE | whisper-ctranslate2 only | — | float16, int8_float16, int8 |
WHISPER_BEAM_SIZE | whisper CLI (sent only if set) | — | 5 |
WHISPER_WORD_TIMESTAMPS | whisper CLI (sent only if set) | off | 1 |
WHISPER_HF_MODEL | HF transformers (opt-in) | — (strategy off) | Xenova/whisper-small |
OPENAI_API_KEY | OpenAI API | — | sk-… |
The default
tinymodel is fast but weak for non-English audio. For Portuguese (or other non-English) sources, install the CLI and setWHISPER_MODEL=small(ormedium) +WHISPER_LANGUAGE=ptfor much better accuracy. AddWHISPER_PROMPTwith a domain glossary (brand/place names) to fix proper nouns. You can also overridemodel/language/initialPromptper call onanalyze_video/get_transcript/analyze_videos— no restart needed.GPU (faster-whisper):
whisper-ctranslate2(pip install -U whisper-ctranslate2) is a drop-in CLI with the same flags plus--device cuda/--compute_type/--beam_size. PointWHISPER_BINat it and setWHISPER_DEVICE=cuda(+ optionallyWHISPER_COMPUTE=float16). These GPU flags are env-gated — they're only passed when set, so plainopenai-whisper(which rejects--compute_type) keeps working when they're unset.Windows note: pip installs
whisper.exeinto the PythonScripts/dir, which is often not on thePATHthat GUI-launched MCP clients inherit. If transcripts come back empty, setWHISPER_BINto the full path ofwhisper.exe.
Frame extraction uses a two-strategy fallback chain — no single dependency is required:
| Strategy | How it works | Speed | Requirements |
|---|---|---|---|
| yt-dlp + ffmpeg (primary) | Downloads video, extracts frames via scene detection | Fast, precise | yt-dlp (pip install yt-dlp) |
| Browser (fallback) | Opens video in headless Chrome, seeks to timestamps, takes screenshots | Slower, no download needed | Chrome or Chromium installed |
The fallback is automatic — if yt-dlp is not available, the server tries browser-based extraction via puppeteer-core. If neither is available, analysis still returns transcript + metadata + comments, just no frames.
After frame extraction, the pipeline automatically applies:
| Step | What it does | Why |
|---|---|---|
| Frame deduplication | Removes near-identical consecutive frames using perceptual hashing (dHash + Hamming distance) | Screencasts often have long static moments — dedup removes redundant frames, saving tokens |
| OCR | Extracts text visible on screen from each frame (via tesseract.js). Each frame is first preprocessed — grayscale + 2× upscale + contrast normalization + sharpen — which materially improves accuracy on stylized overlays (prices, dates, coupons, CTAs). | Captures code, error messages, terminal output, UI text that the transcript doesn't cover |
| Annotated timeline | Merges transcript timestamps + frame timestamps + OCR text into a single chronological view | Gives the AI a unified "what was said, what changed visually, and what text appeared" at each moment |
The OCR step requires tesseract.js (included as a dependency). If it fails to load, analysis continues without OCR — no frames or transcript are lost. OCR preprocessing is on by default; set MCP_OCR_PREPROCESS=0 to OCR the raw frames instead.
OCR always reads the full-resolution frame, not the copy emitted to the client. The two have different jobs: the emitted frame is capped for token cost, while recognition needs every pixel it can get.
Emitted frames are capped at 800 px wide, which suits the common case — talking-head clips, Reels, bug repros — where the subject fills the frame.
It is the wrong size for a dense UI capture: a terminal, dashboard, IDE or spreadsheet recording, where the meaning lives in small text. An unscaled 1920×1080 screen recording lands at 800×450, and a 15 px UI font drops below what a vision model can resolve.
Pass maxWidth per call to keep more (or all) of the source resolution — 0 disables the cap:
Supported on analyze_video, analyze_videos, analyze_moment, get_frames, get_frame_at and get_frame_burst, and on the CLI as --max-width <px>.
Native frames cost several times more context than the default, so raise the cap deliberately — get_frames returns up to 20 frames and analyze_video at detailed up to 60.
| Variable | Applies to | Default | Notes |
|---|---|---|---|
MCP_FRAME_MAX_WIDTH | Emitted frame width, in px | 800 | 0 (or native/full/original) disables the cap. A per-call maxWidth wins over it |
MCP_FRAME_JPEG_QUALITY | Emitted frame JPEG quality | 70 | Raise it when thin glyphs matter; env only, there is no per-call quality parameter. Values outside 1–100 fall back |
MCP_CACHE_DIR | Root for the tessdata cache and the CLI's default --out | per-user cache dir | Absolute paths only (a relative value is ignored). Use it when $HOME is read-only or absent — a hardened container, ProtectHome=, a quota'd home. The published Docker image sets it to /tmp/mcp-video-analyzer-cache so --read-only --tmpfs /tmp works out of the box |
MCP_ALLOW_PRIVATE_URLS | Reaching private/loopback network addresses | off | 1 allows localhost, LAN addresses (192.168.x, 10.x, …), .local names and UNC paths. Off by default — see Network destinations below. Cloud metadata endpoints stay blocked either way |
A value either variable can't use — 1e3, 1920px, a quality of 150 — is rejected with a one-time warning on stderr and the default applies. It is not silently accepted: the whole point of the setting is to escape a downscale that otherwise looks like a normal result.
Prefer the per-call parameter: the server starts once per session, so an environment variable cannot differ between an overview of a YouTube clip and a close read of a screen recording. The width a call actually uses is part of the cache and sidecar key, so analyzing the same video at 800 px and then at maxWidth: 0 re-runs the pipeline instead of returning the first result twice.
For live web debugging alongside video analysis, pair this server with the Chrome DevTools MCP:
When to use each:
| Scenario | Tool |
|---|---|
| Bug report recorded as a Loom video | mcp-video-analyzer — extract transcript, frames, and error text from the recording |
| Live debugging a web page | Chrome DevTools MCP — inspect DOM, console, network, take screenshots |
| Video shows UI issue, need to reproduce it | Use both: analyze the video first, then open the page in Chrome DevTools to reproduce |
The two MCPs complement each other: video analyzer understands recorded content, DevTools interacts with live pages.
The examples/loom-demo/ folder contains real outputs from analyzing a public Loom video (Boost In-App Demo Video, 2:55).
| File | What it shows |
|---|---|
metadata.json | Title, duration, platform |
transcript.json | 42 timestamped entries with speaker IDs |
timeline.json | Unified chronological view (transcript + frames merged) |
moment-transcript-0m30s-0m45s.json | Filtered transcript for analyze_moment (0:30–0:45) |
full-analysis.json | Complete analyze_video output |
Frame images (19 total in examples/loom-demo/frames/):
scene_*.jpg — scene-change detection (key visual transitions)dense_*.jpg — 1fps dense sampling (every 10th frame saved as sample)burst_*.jpg — burst extraction for moment analysis (0:30–0:45)Regenerate after changes:
npx tsx examples/generate.ts— requires yt-dlp + network access.
MIT