yt-dlp search, download, metadata, delivery, and Plex workflows over MCP and CLI.
Omnimodal MCP server converting ComfyUI workflows into tools for text, image, sound, and video generation with web UI.
Key Features: Full-modal support for text, image, sound, and video generation Dual execution modes: local ComfyUI and RunningHub cloud service Zero-code workflow-to-MCP tool conversion Convert HTML slide carousels to PNG, WebP, PDF, or PPTX with themed templates and Puppeteer rendering.
Key Features: Render slides to PNG, WebP, PDF, and PPTX Supports 7 distinct slide themes with design system styles Dimension-aware reflow for portrait and landscape layouts Locally process video or audio into timestamped transcripts, speakers, scenes, chapters, OCR, and searchable timeline data for AI agents.
Key Features: Word-level transcripts in approximately 100 languages Speaker diarization with local processing Scene detection and per-scene keyframes Let any LLM actually watch a video: scene-aware keyframes plus a timestamped transcript, local.
Transforms vague prompts into platform-optimized prompts for 60+ AI platforms across 7 categories.
Key Features: Context-aware prompt optimization for 60+ platforms Supports 7 AI categories: image, video, voice, music, code, chat, document Custom platform registration and management Hosted media-generation MCP server for images, video, audio, transcription, and multi-step workflows via OAuth.
Key Features: Streamable HTTP JSON-RPC 2.0 API OAuth 2.1 authentication with dynamic client registration Multi-modal media generation (image, video, audio) Analyzes videos from URLs or local files to extract transcripts, key frames, OCR text, and annotated timelines without authentication.
Key Features: Supports major platforms: YouTube, Vimeo, TikTok, Instagram, X, Twitch, Dailymotion, Facebook, Loom Processes direct video URLs and local video files Extracts transcripts, key frames, OCR text, and metadata MCP server for generating and editing images and videos via OpenRouter or OpenAI-compatible APIs with file saving and previews.
Key Features: Text-to-image generation with multiple variations and aspect ratios Image editing supporting local files, URLs, and compositing Text-to-video and image-to-video generation with resumable jobs Scores AI image/video generation prompts and creatives against proven winning ads before spending generation credits.
Key Features: Scores prompts on a 0-100 scale against 18 models including Veo 3, Sora, Kling, and Midjourney Scores creatives by embedding and nearest neighbor search against 800+ winning ads Provides enhancement briefs for prompt rewriting based on real market lift data Local-first video editing MCP server with typed FFmpeg tools, captions, effects, quality gates, and provenance receipts.
Key Features: Typed FFmpeg-based video editing tools Video Receipts for provenance and auditability Quality gates and release checkpoints Video scene understanding for AI agents via the Primate Vision API.