Local, offline transcription, speakers, keyframes, on-screen text and review of any audio or video.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent — or use 1-click editor setup below.
One-click editor setup isn’t available for this listing yet — we don’t have a confirmed install command, and we’d rather show nothing than point your editor at the wrong package or host. Follow the project’s own setup instructions, linked above.
Eyes and ears for AI agents. Local, offline transcription, keyframes, on-screen text and a pre-publish review of any audio, video or image — as an MCP server, a CLI and a Node library. No Python, no cloud, no API key.
Docs: Tool reference · Design doc — decisions and measurements · Evaluation results · Contributing · Security · Changelog
11 minutes of screencast become 14 contact sheets and 3 KB of text. And it tells you if your API key is visible at 2:50.
Ollos is Galician for eyes.
Contents: Why · Install · Tools · Sources · CLI · Library · How it works · Evaluation · Requirements & performance · Privacy & security · Configuration · Troubleshooting · Known limits · Roadmap
Agents can't hear or watch. Today you either pay a transcription API, install a Python pipeline, or paste frames by hand. Ollos runs Whisper, speaker segmentation, perceptual-hash keyframing and OCR in Node, through ONNX Runtime, on your machine. The file never leaves it.
It was built for one workflow first — reviewing a screen recording before publishing — and grew into the general case: meetings, lessons, podcasts, downloaded videos.
Node 20+. npm install brings its own ffmpeg (ffmpeg-static); a system ffmpeg is used if present.
Claude Code
or in the project's .mcp.json (the same JSON works for Claude Desktop's claude_desktop_config.json, Cursor's .cursor/mcp.json and Windsurf's mcp_config.json):
VS Code — .vscode/mcp.json uses servers instead of mcpServers:
Claude Desktop reads ~/Library/Application Support/Claude/claude_desktop_config.json on macOS and %APPDATA%\Claude\claude_desktop_config.json on Windows. Environment variables (OLLOS_HOME, OLLOS_YTDLP, …) go in an env object next to args; use absolute paths, ~ is not expanded.
CLI
Models download on first use into ~/.ollos/models. Set OLLOS_OFFLINE=1 afterwards to forbid all network access.
Ten tools, one per distinct contract. Long work never blocks: it returns a jobId you poll.
| Tool | What it does |
|---|---|
ollos_probe | What the file really is: kind, duration, resolution, aspect (and which platforms it fits), codecs, tracks. Detects Zoom recording folders. Instant. |
ollos_transcribe | Whisper transcription with timestamps. Voice-activity gating skips silence; known hallucinations are filtered; vocabulary fixes domain terms. |
ollos_keyframes | The frames that carry information, packed into 3×3 timestamped contact sheets. Works on screen recordings where scene detection sees nothing. |
ollos_read_screen | OCR of on-screen text plus a secret scan: API keys, JWTs, .env lines, private deployment URLs. Always masked. |
ollos_review | Verdict before publishing: loudness vs platform, silences to cut, aspect ratio, secrets on screen. |
ollos_frames | Look at a sheet or a single frame as an image. |
ollos_diarize | Who spoke when: pyannote segmentation + WeSpeaker embeddings + clustering, with an 8-second voice clip per speaker so you can name them by ear. Uses Zoom per-participant tracks directly when present. Experimental — see limits. |
ollos_search | Hybrid BM25 + multilingual-embedding search over everything transcribed and read, fused by reciprocal rank. Returns passages with timestamps, never whole transcripts. |
ollos_job · ollos_cancel | Poll and stop jobs. Jobs live on disk and survive a server restart. |
Every parameter is documented in docs/TOOLS.md (one anchor per tool, e.g. ollos_transcribe); the tool descriptions the agent sees carry the same information.
source accepts a local path, a file:// URL, a Zoom local-recording folder (one audio track per participant), a direct https:// media URL, a video-site URL (YouTube, Instagram, TikTok, Vimeo, X, Loom… through yt-dlp) and a data: URI. URLs are downloaded once into the cache; the download runs inside the job and can be cancelled. Refused: private, loopback and link-local addresses on any redirect hop (OLLOS_ALLOW_PRIVATE=1 to allow), downloads over OLLOS_MAX_DOWNLOAD_MB, media over OLLOS_MAX_DURATION_SEC from any origin.
Tested commands, yt-dlp setup and proxy notes: examples/url-sources.md.
Results are concise by default and point to MCP resources (ollos://jobs/<id>/transcript, /ocr, /report, /sheet/<n>) for the full artifacts, so a 2-hour meeting doesn't flood the context window. Pass format: "detailed" when you want it all.
Add --json for machine output.
More: examples/library.ts (every capability) and examples/library-url.ts (a YouTube URL as the source).
ollos-mcp/core has no MCP dependency: use it from n8n, a script, a Lambda.
Numbers below were measured on an 11:37 screencast (1890×1080, webcam overlay) on a 16-core laptop. They are why the design is what it is.
Transcription. Silero VAD marks speech; Whisper only sees speech (fewer hallucinations, 20–40% less work on meetings). whisper-large-v3-turbo at 1.7× real time got "MCP servers", "n8n", "VS Code" right where whisper-base (5.4×) got all three wrong. The one phonetic miss left ("Cloud Code") is fixed by vocabulary. Two Whisper sessions in parallel measured slower than one (0.6–1.0×), so ASR concurrency is 1 and speed comes from VAD and from running vision in parallel instead.
Anti-hallucination. Whisper doesn't go quiet on silence — it invents "Obrigado." and "Subtitles by the Amara.org community". Four filters, from production experience shared by the Vexa project: exact blocklist per language, repetition-loop collapse, no-speech gate, impossible speaking rate.
Keyframes. ffmpeg scene detection at 0.3 kept 4 frames in 11 minutes of screencast; mpdecimate removed 0% (the cursor and streaming text change every pixel). A 64-bit perceptual hash (dHash) at Hamming ≥ 6 kept 20% — one frame every 5–8 s — and that is the default. Hard cuts, transcript anchors and a 20-second floor fill the gaps.
OCR. Tesseract on a full 1890-px frame missed an on-screen URL entirely; on a 3× upscaled tile it read it whole at 90% confidence in 2.8 s. So OCR runs per tile, and URL-like text is re-joined when OCR splits it ("up. railway .app").
Secrets. Three signals, because OCR garbles the secret more often than the words around it. On a real "API Key Created" modal the plain JWT regex missed (OCR read eyJ as eyl), the entropy detector caught the 157-char token, and the UI context read at 66%. With an OCR-tolerant JWT pattern, native-resolution frames and a centre tile, the end-to-end run now reports it as high · jwt · near "API Key" → block. The first version also produced 58 false positives by running the entropy test on whitespace-stripped text; that is a regression test now. Values are always masked; the tool that warns about a leak must not be the leak.
No reviews yet — be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/ollos-mcp)<a href="https://allmcps.com/mcp/ollos-mcp"><img src="https://allmcps.com/api/badge/ollos-mcp?style=directory" alt="Ollos MCP on AllMCPs" /></a>