The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the MCP Modal Server listing page.
An MCP server for managing Modal — apps, containers, volumes, and secrets — and for deploying & running Modal apps directly from Claude Code and other MCP clients.
Every tool shells out to your local modal CLI, so it operates against whatever Modal profile and credentials are configured on your machine. There are no extra tokens to manage.
The server is published on PyPI as mcp-modal. No manual install is needed — the recommended way to run it is with uvx, which fetches and launches it on demand. Just point your MCP client at the command below (see Configuration).
Every version is also tagged and published on the
Releases page, with release notes and
the same .whl / .tar.gz that PyPI serves attached — useful for pinning, air-gapped
installs, or reading what changed between two versions.
This server uses your local Modal credentials. If you haven't authenticated yet, run:
This opens a browser to log in and stores a token in ~/.modal.toml. Already logged in elsewhere? Check with modal profile current.
Add the server to Claude Code with the claude mcp CLI:
Or add it to a .mcp.json file in your project root, which is the better option for a team
— everyone who opens the repo gets the same configuration:
@latest, and when to pin insteaduvx caches the environment it builds on the first run and does not check PyPI again:
"uvx will use the latest available version of the requested tool on the first invocation. After that, uvx will use the cached version of the tool unless a different version is requested, the cache is pruned, or the cache is refreshed." — uv docs
So a plain uvx mcp-modal means latest at install time, frozen forever after — restarting
the client or rebooting changes nothing, because the cache lives on disk. Different people
end up on different versions depending on when they first ran it, with no warning.
mcp-modal@latest re-resolves on every launch, so a restart picks up new releases.
Costs one network round-trip at startup. Use it while the tool surface is still moving.mcp-modal@0.4.0 (an explicit version) is reproducible and upgrades become a
deliberate one-line change. Use it once you want stability, or for a wider audience.To move a machine that is already stuck on an old cached build, switching it to either form
above is enough — requesting a version invalidates the cache. Otherwise
uv cache clean mcp-modal forces a refresh.
uv (provides uvx)modal setup) — 1.5 is
where modal billing summary/rates landed and where the billing report switched to
snake_case columns; the cost tool reads both spellings but needs 1.5 for those two viewsuv for dependency managementmodal must be installed in that project's virtual environmentThis server shells out to your local modal CLI using whatever credentials are in
~/.modal.toml. A few tools are powerful by design — if the MCP client driving the server
is ever prompt-injected (for example by malicious text inside logs it fetched), these are
the escalation paths and should stay behind your client's tool-approval prompts rather
than being auto-approved:
deploy_modal_app / run_modal_app — execute arbitrary local Python on the host
(modal deploy imports the app file; uv run resolves and installs the target project's
dependencies).modal_volume_files with action="put" — can read any local file (e.g. ~/.ssh/id_rsa,
~/.modal.toml) and upload it to a cloud volume (a data-exfiltration primitive).modal_volume_files with action="get" and force=True — can overwrite any local path
(e.g. ~/.zshrc or a shell profile, a persistence primitive).manage_modal_container with action="exec" — runs arbitrary commands inside a
container, by design.Every tool declares MCP tool annotations,
so a client can distinguish the four read-only tools (list_modal_resources,
get_modal_logs, search_modal_logs, analyze_modal_costs — all readOnlyHint: true)
from the eight that change remote state or start compute. Six of those eight are
destructiveHint: true; the exceptions are run_modal_app and inspect_modal_secret,
which start compute without removing or overwriting anything. Auto-approve the reads; keep
the rest behind a prompt.
To contain the two filesystem-touching volume tools, set the
MCP_MODAL_ALLOWED_LOCAL_PATHS environment variable to an
os.pathsep-separated list of
directories (: on macOS/Linux). When it is set, modal_volume_files is refused for any
local path — local_path on action="put", the destination on action="get" — unless the
resolved path, after expanding ~ and collapsing ../symlinks, falls inside one of those
roots. The download target "-" (return contents instead of writing a file) is exempt
because nothing is written to disk.
When the variable is unset (the default) there is no restriction, so existing setups are unaffected. Configure it in your MCP client, e.g.:
All tools also pass user-supplied names/paths after a -- end-of-options separator, so a
value beginning with - is always treated as data, never as a modal CLI flag. Secret
values handed to manage_modal_secret are redacted from the echoed command, logs, and any
error output.
12 tools. Related operations are grouped behind an action/resource argument rather than
split one-per-CLI-subcommand: every tool schema is loaded into the model's context for the
whole session, so a smaller surface leaves more room for your actual work (and gives the
model fewer near-identical tools to choose between).
Tools that talk to environment-scoped resources take an optional env argument to target a
specific Modal environment; if omitted, they
use the profile's default (or MODAL_ENVIRONMENT). The exception is manage_modal_container
and container logs — a container ID is globally unique and the CLI accepts no environment
there.
List Modal Resources (list_modal_resources) — one lookup for the whole account.
resource (required), name, path (default /), envresource values:
| value | returns | name means |
|---|---|---|
apps | deployed/running/recently-stopped apps | — |
app_history | one app's deployment versions (for rollback) | app name/ID |
containers | running containers (ta-...) | app ID to filter by |
volumes | named volumes | — |
volume_files | files inside a volume (with path) | volume name |
secrets | secret names (values are never exposed) | — |
environments | valid env values for this workspace | — |
profile | active profile + all profiles | — |
volume_files sets empty: true with a message when a listing genuinely returns
nothing, so an empty directory is distinguishable from a wrong path.omitted_items giving the number dropped.Get Modal Logs (get_modal_logs) — fetch or stream logs for an app or a container.
identifier (required), target (auto/app/container, default
auto — anything starting ta- is a container), timeout_seconds (default 30),
env, since, until, tail, source (stdout/stderr/system), timestamps,
followsince without tail fetches every entry in the range; pass until as well (max
range 35 days, tail max 20,000) to keep a busy app's output bounded.follow=True, logs stream until the app/container stops or timeout_seconds is
reached, returning a snapshot with truncated: true.Search Modal Logs (search_modal_logs) — grep logs and get each hit with the
surrounding lines, built for "where did it go wrong?" debugging. Logs are fetched once
and searched locally, so you get context, regex, case control, and exact match counts.
identifier (required), pattern (required), target (default auto),
regex, case_sensitive, context_lines (default 3), max_matches (default 50),
since, until, tail (defaults to the last 1000 entries), source,
exclude (drop noise lines before searching, e.g. "queue put failed"),
prefilter, timestamps (default true), timeout_seconds, envsince on its own fetches everything from then
until now — hundreds of KB per hour on a chatty app, which the 30s fetch cuts off
(logs_truncated: true) and the output budget trims. since and until around the
minute you care about is the fix, and is usually kilobytes.prefilter=True pushes pattern down to Modal as a server-side substring filter
(modal app logs --search), so non-matching lines are never fetched — the lever for
logs too large to drain. Requires regex=False, and context lines then show only
other matches, so use it to locate the window and re-query it with prefilter=False.match_count and matches: timestamped, line-numbered context blocks where
matched lines are prefixed with >, e.g. > 8: 2026-06-04T... ValueError: bad input.
The whole fetched log is always searched, so match_count stays exact even when fewer
blocks are returned. returned is how many matches came back (adjacent matches merge
into one block, counted by returned_blocks). Reports excluded_lines when exclude
is used.tail over 20,000) comes back
as success: false with Modal's own message, not a bare exit code.get_modal_logs.Deploy Modal App (deploy_modal_app)
modal deploy). Deployed web endpoints persist, so any links in
the output are live and shareable (returned in urls).absolute_path_to_app (required), env, name, tag,
strategy (rolling/recreate), stream_logsuv with modal installed in its virtualenv.Run Modal App (run_modal_app)
modal run).absolute_path_to_app (required), function_name, env, detach,
timeout_seconds (default 120)truncated: true if the run is still going at the timeout.
Pass detach=True to keep long jobs alive on Modal past the timeout.Why no
modal servetool?modal serveonly keeps its endpoints alive while the blocking process runs — an MCP tool that returns would tear them down immediately, handing back a dead URL. Usedeploy_modal_appfor a persistent, shareable endpoint.
Manage Modal App (manage_modal_app) — action is stop (shut the app down and
terminate its containers) or rollback (redeploy a previous version).
action (required), app_identifier (required), version (rollback
only — defaults to the immediately preceding version), envManage Modal Container (manage_modal_container) — action is exec (run a command
inside a running container, modal container exec --no-pty) or stop (terminate it).
action (required), container_id (required), command (exec only —
a list of args, e.g. ["python", "-c", "print('hi')"]), timeout_seconds (default 60)Manage Modal Volume (manage_modal_volume) — action is create, delete
(the volume and all its data, irreversible), or rename.
action (required), volume_name (required), new_name (rename only), envModal Volume Files (modal_volume_files) — write operations on a volume's files:
action is put (upload), get (download), cp (copy inside the volume), or rm.
action (required), volume_name (required), local_path, remote_path,
paths (for cp: sources then destination), recursive, force, envaction="get" with local_path="-" returns the file contents instead of writing a file.list_modal_resources(resource="volume_files").Manage Modal Secret (manage_modal_secret) — action is create or delete.
action (required), secret_name (required), key_values (dict),
from_dotenv (path), from_json (path), force, env. Creating requires at least
one of key_values, from_dotenv, or from_json.list_modal_resources(resource="secrets").analyze_modal_costs) — read-only. Fetches
modal billing once and aggregates locally, so you get ranked totals and
period-over-period changes instead of hundreds of raw rows.
view (default by_app), period, start, end, resolution
(d/h), timezone, app, environment, top_n (default 10), tag_namesview values:
| value | answers |
|---|---|
by_app | "what is my costliest app?" — apps ranked by spend, with % share |
timeline | "why was Monday expensive?" — cost per interval, plus an explanation that diffs the peak interval against the one before and ranks which apps grew |
by_environment | which environment the money goes to |
by_resource | CPU vs GPU class vs memory vs storage |
summary | billed vs metered cost for a month cycle, with credits/plan adjustments |
rates | current unit prices |
total_cost always covers every row in range, even when groups is cut to
top_n — quote it rather than summing the visible rows.-e), so this reports across all
environments; environment filters the rows afterwards.inspect_modal_secret) — lists the key names inside a
secret, never the values.
secret_name (required), env, image, timeout_seconds (default 300)modal shell --secret <name> with
compgen -e (a bash builtin that prints exported variable names only — no value is
ever printed, even inside the container), then subtracts the variables the image and
Modal runtime set anyway — 23 known names plus anything under six prefixes
(MODAL_, PYTHON, PIP_, NVIDIA_, CUDA_, LD_LIBRARY_PATH), which also
covers the MODAL_TOKEN_* credentials that live in every container.list_modal_resources(resource="secrets") to see which secrets exist and
reach for this only when you need to know what is inside one.keys, plus the unfiltered all_env_names so a key that looks like a
runtime variable is still visible rather than silently dropped.image to use Modal's default (built to match the server's Python — the most
reliable choice). Pass one, e.g. python:3.12-slim, if your workspace's image
builder rejects that Python version.The server also ships four MCP prompts — multi-step workflows your client can invoke
directly (in Claude Code they appear as /mcp__mcp-modal__<name>). Prompts are fetched on
demand, so unlike tools they cost nothing in per-session context:
debug_modal_app (app_name, optional symptom) — an ordered triage routine: check
the app is up, search logs for tracebacks with context, narrow the window instead of
widening it when a log fetch comes back truncated, fall back to the log tail, check
whether sibling apps were hit in the same window, inspect containers, then compare
against deployment history and consider a rollback.deploy_and_verify (absolute_path_to_app, optional env) — confirm the target
workspace, deploy, report the live URLs, then verify the app is healthy instead of
assuming it.review_modal_account (optional env) — a read-only inventory that flags idle apps,
unexplained running containers, and orphaned volumes/secrets, naming the exact call that
would clean each one up without running it.investigate_modal_costs (optional period, app) — traces a spend increase from
the daily timeline down to the peak hour, the resource class, and the deploy or
still-running container behind it.Log, run, and exec output is capped before it is returned, so one chatty app can't flood
your context window. The default budget is 40,000 characters per text field (roughly 10k
tokens); when a field is trimmed the result sets output_capped: true and the text carries
a marker naming how much was dropped. A capped field keeps its head and its tail, so a
startup banner and the traceback at the end both survive.
Searching is never capped before the fact: search_modal_logs greps the whole fetched log
and only limits how many context blocks come back, so match_count is always exact.
Raising timeout_seconds or the budget is rarely the right answer to a truncated log
search — fetching less is. Bound the window with since and until, filter with
source/exclude, or set prefilter=True to drop non-matching lines inside Modal.
Set MCP_MODAL_MAX_OUTPUT_CHARS to raise or lower the budget, or to 0 to disable capping
entirely:
All tools return responses in a standardized format, with slight variations depending on the operation type:
This project is licensed under the MIT License - see the LICENSE file for details.