Multi-provider media generation β images, video, audio, and transcription via a unified interface
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
π‘ Paste the JSON block into your client's configuration file under mcpServers, then restart the application.
Multi-provider media generation MCP server. Generate images, videos, audio, and transcriptions from text prompts using OpenAI, xAI, Gemini, ElevenLabs, and BFL (FLUX) through a single unified interface.
Set the API key for at least one provider. Most users only need one β add more to access additional providers.
Using a different editor? See setup instructions for Claude Desktop, Cursor, VS Code, Windsurf, and Cline.
| Variable | Required | Description |
|---|---|---|
OPENAI_API_KEY | At least one provider key | OpenAI API key β enables image, video, audio generation, and transcription via gpt-image-1, sora-2, tts-1, and whisper-1 |
XAI_API_KEY | At least one provider key | xAI API key β enables image and video generation via grok-imagine-image and grok-imagine-video |
GEMINI_API_KEY | At least one provider key | Gemini API key β enables image, video, and audio generation via imagen-4, veo-3.1, and gemini-2.5-flash-preview-tts |
GOOGLE_API_KEY | β | Alias for GEMINI_API_KEY; either name is accepted |
ELEVENLABS_API_KEY | At least one provider key | ElevenLabs API key β enables audio generation (TTS, sound effects) and transcription via Flash v2.5 and Scribe v1 |
BFL_API_KEY | At least one provider key | BFL API key β enables image generation and editing via FLUX Pro 1.1 and FLUX Kontext |
MEDIA_OUTPUT_DIR | No | Directory for saved media files. Defaults to the current working directory |
generate_imageGenerate an image from a text prompt.
| Parameter | Type | Required | Description |
|---|---|---|---|
prompt | string | Yes | Text description of the image to generate |
provider | string | No | Provider to use: openai, xai, google, bfl. Auto-selects if omitted |
aspectRatio | string | No | Aspect ratio: 1:1, 16:9, 9:16, 4:3, 3:4 |
quality | string | No | Quality level: low, standard, high |
outputDirectory | string | No | Directory to save the generated file. Absolute or relative path. Defaults to MEDIA_OUTPUT_DIR or cwd |
providerOptions | object | No | Provider-specific parameters passed through directly |
generate_videoGenerate a video from a text prompt. Video generation is asynchronous and may take several minutes.
| Parameter | Type | Required | Description |
|---|---|---|---|
prompt | string | Yes | Text description of the video to generate |
provider | string | No | Provider to use: openai, xai, google. Auto-selects if omitted |
duration | number | No | Video duration in seconds (provider limits apply) |
aspectRatio | string | No | Aspect ratio: 16:9, 9:16, 1:1 |
resolution | string | No | Resolution: 480p, 720p, 1080p |
outputDirectory | string | No | Directory to save the generated file. Absolute or relative path. Defaults to MEDIA_OUTPUT_DIR or cwd |
providerOptions | object | No | Provider-specific parameters passed through directly |
generate_audioGenerate audio from text. Supports text-to-speech and sound effects. Audio generation is synchronous.
| Parameter | Type | Required | Description |
|---|---|---|---|
text | string | Yes | Text to convert to speech, or a description of the sound effect to generate |
provider | string | No | Provider to use: openai, google, elevenlabs. Auto-selects if omitted |
voice | string | No | Voice name (provider-specific). OpenAI: alloy, ash, coral, echo, fable, nova, onyx, sage, shimmer. Google: Kore, Charon, Fenrir, Aoede, Puck, etc. ElevenLabs: voice ID |
speed | number | No | Speech speed multiplier (OpenAI only): 0.25 to 4.0 |
format | string | No | Output format (OpenAI only): mp3, opus, aac, flac, wav, pcm |
outputDirectory | string | No | Directory to save the generated file. Absolute or relative path. Defaults to MEDIA_OUTPUT_DIR or cwd |
providerOptions | object | No | Provider-specific parameters passed through directly. ElevenLabs: set mode: "sound-effect" for sound effects, model for TTS model selection |
transcribe_audioTranscribe audio to text (speech-to-text).
| Parameter | Type | Required | Description |
|---|---|---|---|
audioPath | string | Yes | Absolute path to the audio file to transcribe |
provider | string | No | Provider to use: openai, elevenlabs. Auto-selects if omitted |
language | string | No | Language code (e.g., en, fr, es) to hint the transcription language |
providerOptions | object | No | Provider-specific parameters passed through directly |
list_providersList all configured media generation providers and their capabilities. Takes no parameters.
| Provider | Image | Image Editing | Video | Audio | Transcription | Key Models |
|---|---|---|---|---|---|---|
| OpenAI | β | β | β | β | β | gpt-image-1, sora-2, tts-1, whisper-1 |
| xAI | β | β | β | β | β | grok-imagine-image, grok-imagine-video |
| Gemini | β | β | β | β | β | imagen-4, veo-3.1, gemini-2.5-flash-preview-tts |
| ElevenLabs | β | β | β | β | β | eleven_flash_v2_5, scribe_v1 |
| BFL | β | β | β | β | β | flux-pro-1.1, flux-kontext-pro |
| Provider | 1:1 | 16:9 | 9:16 | 4:3 | 3:4 |
|---|---|---|---|---|---|
| OpenAI | β | β | β | β | β |
| xAI | β | β | β | β | β |
| Gemini | β | β | β | β | β |
| BFL | β | β | β | β | β |
| Provider | 16:9 | 9:16 | 1:1 | 480p | 720p | 1080p |
|---|---|---|---|---|---|---|
| OpenAI | β | β | β | β | β | β |
| xAI | β | β | β | β | β | β |
| Gemini | β | β | β | β | β | β |
| Provider | mp3 | opus | aac | flac | wav | pcm |
|---|---|---|---|---|---|---|
| OpenAI | β | β | β | β | β | β |
| Gemini | β | β | β | β | β | β |
| ElevenLabs | β | β | β | β | β | β |
Set at least one of OPENAI_API_KEY, XAI_API_KEY, GEMINI_API_KEY, ELEVENLABS_API_KEY, or BFL_API_KEY in the MCP server's env block.
Each provider supports different media types (see Provider Capabilities). If you specify a provider that isn't configured (no API key) or doesn't support the requested media type, you'll receive an error. Omit the provider parameter to auto-select from configured providers.
Video generation polls for up to 10 minutes. If your video hasn't completed in that window, the request will fail with a timeout error. Try a shorter duration or a simpler prompt.
This indicates the xAI API returned an empty response. Check that your XAI_API_KEY is valid and that your prompt does not violate xAI content policies.
Verify your GEMINI_API_KEY has the Generative Language API enabled in Google Cloud Console.
Replace OPENAI_API_KEY with your provider of choice (XAI_API_KEY, GEMINI_API_KEY, ELEVENLABS_API_KEY, BFL_API_KEY). You can set multiple keys to enable multiple providers.
Add to ~/Library/Application Support/Claude/claude_desktop_config.json:
Add to .cursor/mcp.json in your project root (or ~/.cursor/mcp.json globally):
Add to .vscode/mcp.json in your project root:
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/multimodal)<a href="https://allmcps.com/mcp/multimodal"><img src="https://allmcps.com/api/badge/multimodal?style=directory" alt="Multimodal on AllMCPs" /></a>