Self-hosted MCP server: video lipsync, face restore, and ffmpeg ops (trim, concat, transcode).
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent — or use 1-click editor setup below.
💡 Paste the JSON block into your client's configuration file under mcpServers, then restart the application.
Video toolkit. One port. Zero cloud. Lipsync, face restore, ffmpeg. Fire-and-forget async jobs. Webhooks. Spec-first OpenAPI; typed Go + Python clients generated from the same spec.
The video sibling of audiolla (audio) and talkies (speech). Same wire format, same async-job model, same bind-mount-/data story, same Makefile shape, same :latest + :latest-cuda split, same opt-in non-commercial gate.
POST a JSON body. Get a video back. Drive it from curl, shell scripts, the generated Go/Python clients, or point an LLM agent at the MCP endpoint.
No account. No subscription. docker run and you're done.
| 👄 Lipsync | LatentSync 1.5 (ByteDance, Apache-2.0, default on CUDA) + Wav2Lip / Wav2Lip-GAN (Rudrabha, fast/low-VRAM, behind FLICKIES_ENABLE_NONCOMMERCIAL=1) |
| 🧹 Face restore | GFPGAN v1.4 (TencentARC, Apache-2.0) — chains after Wav2Lip to fix the soft 96×96 mouth crop, or use standalone |
| ⚙️ ffmpeg ops | Trim · concat · transcode (incl. gif + fps + codec change) · scale · mux audio · extract audio · thumbnail grid — pure ffmpeg, CPU |
| 📋 Info | ffprobe metadata at /v1/video/info — duration, codec, fps, dimensions, bitrate |
| 🔗 MCP server | All endpoints exposed as MCP tools so function-calling LLMs can drive the pipeline |
| 📜 Spec-first | openapi.yaml is the single source of truth — server-side Pydantic, Go client, and Python client all regenerated from one file |
| 🐳 Hot-swap eviction + idle unload | One GPU pool. Different model requested → current model evicted. Idle longer than FLICKIES_IDLE_UNLOAD_SECS (default 600s) → unloaded by the sweeper. |
CUDA image at psyb0t/flickies:latest-cuda runs every engine at usable speed. CPU image runs all ffmpeg ops (trim/concat/transcode incl. gif/scale/mux/extract/thumbnail-grid/info) + Wav2Lip-CPU (~44s for a 3s clip; OK for short ones). GFPGAN + LatentSync 1.5 are CUDA-only — CPU image refuses to load them.
Weights live in the standard HuggingFace cache layout under /data/hf/hub/models--<org>--<name>/{blobs,snapshots,refs}/… — content-addressed blobs, snapshot-named symlinks, reusable by any other HF-aware tool sharing the bind mount (not just flickies). Sources:
| engine | HF repo |
|---|---|
wav2lip / wav2lip-gan | Nekochu/Wav2Lip |
| S3FD detector | ByteDance/LatentSync-1.5 (bundled in auxiliary/) |
gfpgan | leonelhs/gfpgan |
latentsync-1.5 | ByteDance/LatentSync-1.5 |
Lazy by default — each engine fetches its repo on first request. Set FLICKIES_ENABLED_ENGINES=wav2lip,gfpgan (or FLICKIES_PREFETCH_ALL=1) to pull at boot before uvicorn starts. FLICKIES_OFFLINE=1 disables auto-download (operators stage the snapshot dir manually).
Bearer token set via env. Any string works:
Unset → auth disabled. /healthz is always probe-exempt.
Structured JSON to both stderr AND a rotating file at FLICKIES_LOG_FILE (default /data/logs/flickies.log, 50 MB × 5 backups). Every line carries time (ISO 8601 UTC sub-ms), level, logger, file, line, func, msg, trace_id, request_id + typed extras.
Inbound X-Request-Id (UUID v4 OR ULID; garbage → server mints fresh) threads onto the logging scope via ContextVar + echoes back on the response. Outbound httpx fetches forward X-Request-Id + X-Trace-Id so the next hop's logs correlate. Sensitive keys (authorization, cookie, *token*, *secret*, hf_*, sk-ant-*) get [REDACTED] automatically at format time.
Default level is INFO; set FLICKIES_LOG_LEVEL=DEBUG for reconstruction-grade tracing: every ffmpeg/ffprobe command + result, each transform's decision (e.g. trim stream_copy vs precise_reencode) + output size, engine inference timing (wall_secs), URL fetch/upload byte counts, and job lifecycle. Logged URLs are stripped of their query string so presigned credentials never reach the logs.
Eleven tools at /v1/mcp via streamable-HTTP JSON-RPC: list_engines, info, lipsync, restore, transcode, trim, concat, scale, mux_audio, extract_audio, thumbnail_grid. Point a function-calling LLM at it (LibreChat, Cursor, Claude desktop with the MCP connector) and it drives the pipeline.
Tested target: RTX 3060 12 GB. Fits LatentSync 1.5 (~8 GB) with headroom. Wav2Lip + GFPGAN chain peaks at ~5 GB. One engine resident at a time — different model request triggers hot-swap eviction.
Wav2Lip variants are trained on LRS2 (non-commercial). The server refuses to load them unless FLICKIES_ENABLE_NONCOMMERCIAL=1 is set in the server env. LatentSync 1.5 (Apache-2.0) is the commercial-safe default — no gate.
| Engine | License | Gate |
|---|---|---|
| LatentSync 1.5 | Apache-2.0 | none |
| Wav2Lip / Wav2Lip-GAN | LRS2 non-commercial | FLICKIES_ENABLE_NONCOMMERCIAL=1 |
| GFPGAN | Apache-2.0 | none |
| ffmpeg / ffprobe (not an engine; standard CPU helper) | LGPL (ffmpeg) | none |
Same pattern as audiolla's MusicGen / matchering gates.
Every request/response shape lives in openapi.yaml. The Pydantic models in src/flickies/schema/_generated.py, the Go client in pkg/clients/go/client.gen.go, and the Python client in pkg/clients/python/flickies-client/ are all generated from that single file.
Never hand-edit generated files. Edit openapi.yaml, run make generate, commit everything together.
The skill works in any agent that reads .agents/skills/, and installs natively in the clients below.
Claude Code prompts for the flickies URL and, if auth is enabled, the token — the token is stored in your OS keychain.
Installed via the marketplace, the skill invokes as $flickies:flickies. Codex also picks the skill up automatically with no install in any repo containing .agents/skills/, where it invokes as plain $flickies.
The skill is published to ClawHub on every release:
For MCP clients that speak local stdio, the @psyb0t/flickies plugin bridges to flickies' /v1/mcp endpoint:
Then set FLICKIES_URL (and FLICKIES_AUTH_TOKEN if the server requires one).
Mounts in aigate at /flickies/ and /flickies-cuda/ behind the same nginx → make run-bg lives. FLICKIES=1 and FLICKIES_CUDA=1 toggle the variants.
WTFPL for flickies itself. Bundled models follow their upstream licenses — review before commercial redistribution.
Factual signals from GitHub, npm, and our automated checks — not a rating.
No reviews yet — be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/flickies)<a href="https://allmcps.com/mcp/flickies"><img src="https://allmcps.com/api/badge/flickies?style=directory" alt="Flickies on AllMCPs" /></a>