The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Vidin listing page.
Turn a video into something a coding agent can actually read.
Claude Code, Codex and friends accept images, not video. So when a bug report arrives as a screen recording, the video is dead weight — someone has to watch it and retype what happened. vidin watches it instead: it converts the recording into a small set of annotated keyframes plus a written timeline, and serves them to agents over MCP.

A 20-second recording in, six frames out: the refund panel, the 500 dialog, the retry toast, the recovery — each boxed and timestamped. (The clip is vidin's synthetic test fixture.)
npx @teunlao/vidin — all it asks of the machine
is ffmpeg; no sharp, no headless browser, no native image library.Hook it into Claude Code:
Or in .mcp.json / ~/.codex/config.toml equivalents:
Or straight from the shell:

Each frame carries a red dashed box around the regions that changed since the previous kept frame, and a bottom bar with the frame number, the timestamp in the original video, the gap since the previous frame, why the frame was kept, and how much of the screen changed.
Four tools, shaped so an agent spends context in the right order:
| tool | what it does |
|---|---|
analyze_video | video → frames on disk; returns the report text and one contact sheet image |
get_frames | full-detail frames by number or time range, capped so a call can't flood the context |
zoom_clip | re-extract a narrow window at a higher frame rate — for the moment two frames don't explain |
probe_video | duration / resolution / fps / audio, cheap |
The intended loop: analyse → look at the sheet → read the timeline → open two or three frames → zoom in if the moment between them is still unclear. Opening everything up front would burn the context before knowing which second matters.
| profile | for |
|---|---|
bug-report | default — someone recorded a screen, narrated a bug, stopped |
dense | short clips where every twitch matters (24 fps analysis, up to 80 frames) |
walkthrough | long recordings; fewer frames, wider anchors |
cheap | smallest footprint, for when the context really is tight (12 frames at 1092px) |
Tuning without leaving the profile: --max-frames, --sensitivity (percent of
the screen that counts as an event, lower = more frames), --fps, --width,
--quality, --no-annotate, --no-sheets.
One screenshot per second turns three minutes into 180 frames of which ~90% are identical, because the person was reading, thinking or moving the mouse. Those duplicates don't cost much — they bury the handful of frames where something actually happened, and an agent has to look at all of them to find out which. Frame selection is a signal-to-noise problem wearing a cost problem's clothes.
vidin is event driven instead. It walks the video at 12 fps in downscaled RGB and keeps a frame only when something actually happened:
Two invariants hold whatever path a frame took:
Decoding streams two frames at a time; the selector additionally keeps the
downscaled buffers of the frames it has kept so it can recompute their deltas
after pruning. That is bounded by the frame budget, not by video length: ~32 MB
at the defaults. All drawing is done by ffmpeg filters, which is why there is
no image library in package.json.
Development runs on Bun; the runtime code itself sticks to
node: APIs, which is what lets the published package run anywhere npx does.
The fixture is a synthetic screen recording with a scripted timeline — a panel,
an error dialog, a toast, a recovery, and two stretches of pointer-only motion.
The suite asserts that each staged event yields exactly one settled frame and
that pointer motion alone yields none. test/regressions.test.ts builds its
own clips and pins every failure a review has found, so a fixed bug stays
fixed.
whisper.cpp or mlx-whisper) → a transcript.md
aligned to the frame timeline.--sensitivity 0.2 or --profile dense catches them, at the cost of more
frames.