The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Clickcast listing page.
Give AI agents visual + structured feedback about live web UIs — and give humans deterministic demo reels while you're at it.
What's new (v0.3.0) — live agent control via a new MCP server, pixel-level
clickcast diff, accessibility semantics fused with the pixel-grid overlay, an official CI GitHub Action, day-one Homebrew/apt packaging, and a self-healing first run. SeeCHANGELOG.mdfor the full notes.

See
docs/ONE_PAGE_NAVIGATION_ORDER_TIPS.mdfor the nine principles behind why this reel reads as legibly as it does — and the scenario template you can copy for your own reels.
clickcast drives a real browser through a website and hands back two things:
docs/feedback-schema.md.Point it at a URL and it will auto-discover the interactive elements and build a tour for you, or hand it a small YAML scenario for a scripted, repeatable walkthrough.
One line, ready to run immediately after:
Broken down — pip install clickcast (requires Python ≥ 3.10) gets you the
CLI; clickcast install downloads Chromium (~one-time, ~180MB — kept out of
the pip package itself since it's versioned independently and every project
doesn't need every engine); --with-deps also pulls the system libraries
Chromium needs (Linux only, may prompt for sudo); clickcast doctor
confirms everything above actually worked.
Forgot the second step? Any command that needs a browser (auto, run,
shot, elements, mcp) detects a missing engine itself and offers to
install it right then instead of failing — say yes once and it retries
automatically. That prompt only fires in an interactive terminal; CI/scripted
runs fail fast with the exact fix command instead of hanging on stdin.
.deb)Both native packages skip the Chromium download (~180MB, versioned
independently of clickcast) and never install a second ffmpeg (clickcast
already bundles one via imageio[ffmpeg]) -- run clickcast install --with-deps chromium once after either install path. Full design rationale,
what's live today vs. what needs one-time bootstrapping, and the exact
bootstrap steps: docs/packaging/homebrew.md,
docs/packaging/apt.md.
clickcast + clickcast-mcp)Two npm packages -- thin Node wrappers whose postinstall provisions an
isolated Python venv and pip-installs the pinned PyPI clickcast, since
there's no way to ship clickcast's Playwright/Pillow/ffmpeg runtime as pure
JS. Built specifically because the MCP ecosystem's install pattern is
npx <package>, not pip install:
Both are published on the npm registry. To run against your working copy instead of the published version:
Full design rationale
(including why the shared provisioning code is a vendored copy rather than
a file: dependency), what's live today vs. what needs bootstrapping, and
the exact bootstrap steps: docs/packaging/npm.md.
Paste the block below into your coding agent (Claude Code, Cursor, Copilot Chat, Codex, etc.) to teach it clickcast in one message. The agent will install the tool, verify the environment, generate a visual + machine-readable report for your project, and know how to gate CI on the results.
Once the agent has this, ask it something concrete like "run clickcast auto against http://localhost:3000 and tell me which clicks had DOM reactions" — it now has everything it needs.
Everything above is batch mode: record a whole tour, then read back a GIF + sidecar. clickcast mcp is the live counterpart — an MCP server that drives one action at a time (goto/click/type/scroll/...) and hands back clickcast's richer per-call payload (annotated frame, page_state, grid coordinates, an enumerated error_code) instead of a bare screenshot, so an agent can react before deciding the next step.
Reach for mcp when an agent needs to explore and react live; reach for auto/run for a repeatable, one-shot artifact (CI, docs, release notes). Full tool reference, client config for Claude Code / Claude Desktop, and the schema doc: docs/mcp.md · docs/mcp-tool-schema.md.
Produces two files:
tour.gif — the reeltour.gif.json — the AI-consumable sidecar (schema_version: 1, spec at docs/feedback-schema.md)For a walkthrough of how an LLM agent consumes both, see docs/ai-integration.md.
| Mode | Command | When |
|---|---|---|
| Auto | clickcast auto <url> | Quick tour of a site; you don't care about the exact script. |
| Scenario | clickcast run tour.yml | Precise, repeatable walkthrough. Docs, release notes, CI. |
| Shot | clickcast shot <url> | One screenshot, viewport or full-page. |
All three are deterministic, headless-by-default, and CI-friendly.
auto <url>Discover interactive elements and record a click-tour.
| Flag | Default | Notes |
|---|---|---|
--out PATH | reel.gif | Extension picks the format. |
--max-steps N / -N | 10 | Cap on discovered elements to click. |
--dwell SEC | 1.0 | Hold time after each action. |
--initial-wait SEC | 2.0 | Post-networkidle hold to let SPAs hydrate. |
--viewport WxH | 1280x800 | |
--device NAME | – | Playwright preset (e.g. "iPhone 15", "Pixel 8"). |
--engine E | chromium | chromium / firefox / webkit. |
--headful | off | Show a real browser window. |
--lang LOCALE | – | e.g. en-US. |
--dark | off | Emulate prefers-color-scheme: dark. |
--fps N | 12 | |
--format F | – | Override extension-derived format. |
--quality 1..30 | 8 | Lower = better (higher fidelity, bigger file). |
--loop N | 0 | 0 = infinite. |
--no-sidecar | off | Skip the JSON. |
-v / --verbose | – | Repeatable. |
run <scenario.yml>Execute a YAML scenario. See Scenario format.
Flags: --out, --format, --headful, --slowmo MS, --url URL (retarget the first goto step — see below), --var key=value (repeatable — substitute {{ key }} inside the scenario), --no-sidecar.
CLI flags override the scenario's meta: block.
Point an existing scenario at a different environment with --url — no YAML edits, no {{ URL }} templating:
--url rewrites the first goto step's URL and wins over --var URL=.... Only the first goto is touched — later goto steps are usually intra-app navigation from the entry point, so they stay put.
shot <url>Single screenshot.
Flags: --out, --full-page, --wait (load / domcontentloaded / networkidle / a selector / a number of seconds), --viewport, --device, --engine, --dark.
init [path]Scaffold a starter YAML scenario. --from-auto runs discovery once and seeds the file with the top-scoring click steps.
Flags: --url, --name, --out, --from-auto, --force.
elements <url>Dump the discovered interactive elements — useful for authoring selectors.
Each entry additionally carries an accessibility block (Playwright's own
role / accessible name / interactive state, fused with the pixel-grid
overlay's grid_cell when --grid is on) — see
docs/feedback-schema.md.
| Flag | Default | Notes |
|---|---|---|
--limit N | 20 | Cap on returned elements. |
--json | off | Emit machine-readable JSON on stdout. |
--viewport WxH | 1280x800 | |
--device NAME | – | Playwright preset (e.g. "iPhone 15", "Pixel 8"). |
--engine E | chromium | chromium / firefox / webkit. |
--headful | off | Show a real browser window. |
--lang LOCALE | – | e.g. en-US. |
--dark | off | Emulate prefers-color-scheme: dark. |
--slowmo MS | 0 | Delay each Playwright op by N ms. |
-v / --verbose | – | Repeatable. |
--grid | off | Populate each element's accessibility.grid_cell (#196). |
--grid-pitch N | 100 | Major-line spacing in px, used for grid_cell. |
--grid-color HEX | #FFFFFF33 | Unused by elements beyond validation — kept symmetric with auto/run/shot. |
--grid-style full|ruler | full | Unused by elements beyond validation — kept symmetric with auto/run/shot. |
doctorCheck Python version, playwright, engine binaries, ffmpeg, config path.
configRead / write persistent defaults.
Set values land in the user TOML at clickcast config path. See Configuration for precedence.
install [engines…]Wrapper over playwright install. Default engine: chromium.
mcpStart a stdio MCP server for live agent control — see Live agent control (MCP) above and docs/mcp.md. Requires pip install 'clickcast[mcp]'.
Flags: --engine, --viewport, --device, --headful, --lang, --dark, --grid/--grid-pitch/--grid-color/--grid-style — all defaults for start_session when the connecting agent doesn't override them.
A scenario is plain YAML: a meta: block and a list of steps:. Full worked examples: docs/scenarios/.
| Action | Example | Notes |
|---|---|---|
goto | goto: https://… | Navigate. Pair with wait. |
click | click: "text=Compare" | CSS, text=…, or role=… selectors — Playwright syntax. Also accepts click: { selector: ..., wait: networkidle } — wait (same shape as goto's) blocks after the click; pair it with a target that triggers client-side/SPA navigation. |
dblclick | dblclick: ".cell" | Also accepts wait, same as click. |
hover | hover: ".menu" | Reveals CSS :hover state. |
type | type: { into: "#q", text: "Japan", delay: 40 } | delay is per-char ms. |
press | press: "Enter" | Or press: { key: "Ctrl+A", selector: "#in" }. |
select | select: { in: "#m", value: "GDP" } | in: in YAML → into internally. |
scroll | scroll: { to: "footer" } or scroll: { by: 600 } | Element or pixel scroll. |
wait | wait: 1.5 or wait: networkidle or wait: ".map-loaded" | Number = seconds, string = load-state or selector. |
screenshot | screenshot: { full_page: true } | Force a frame capture. |
Every step also accepts label, dwell, optional: true (don't fail the run if the selector is missing — sidecar records status: "skipped"), and repeat: N.
Variable substitution: {{ key }} inside any string, injected via --var key=value.
Fluent, chainable — every builder returns self:
Async variant for callers already inside a running event loop:
Discovery only, no reel:
Skip the sidecar with save(..., no_sidecar=True).
Building a static site and reeling it locally before pushing? Use Reel.serve_dir (or the standalone serve_directory helper) as a context manager — it starts a threaded HTTP server on a free port, yields the base URL, and tears the server down on exit. No more python3 -m http.server 8091 & invocations that leak past your shell session and collide with the next iteration.
Defaults are safe for dev iteration: loopback-only bind (127.0.0.1), OS-picked free port, ThreadingHTTPServer so parallel browser requests don't queue. Override any of them explicitly:
bind="0.0.0.0" exposes the server to your LAN — opt-in, not the default.
Every recording run writes <out>.json alongside the media file.
Consumers that don't want to import clickcast can parse the JSON directly against the schema at src/clickcast/feedback/schema/v1.json. A standalone reference implementation lives at tests/consumer/read_sidecar.py.
See docs/ai-integration.md for the two-line agent-integration example and docs/feedback-schema.md for the full field-by-field walkthrough.
On GitHub Actions? .github/actions/clickcast
packages this recipe (plus the visual diff gate below, browser caching,
and a PR comment with the reel + a summary table) as a reusable Action —
see docs/ci/README.md. This section stays the
canonical recipe for every other CI platform, and for anyone who'd rather
not depend on a third-party Action for two lines of shell.
Every reel writes a JSON sidecar, but raw sidecars carry timestamps, frame
filenames, and query-string tokens — none of which are stable across runs.
For a proper CI regression gate, use clickcast assertions (or
Reel.assertions()) to distill the sidecar down to the shape that
actually matters: step count, per-step action / label / status, and the
per-step error counters.
The distilled shape is byte-identical across runs of the same scenario
against the same URL (schema: docs/assertions-schema/v1.json).
Diff it against a committed baseline; non-zero exit on drift.
Bootstrap the baseline once:
Then in CI (2 lines):
Exit 0 means the target UI produced the same step ordering, statuses, and
error-signal counts as when the baseline was captured; anything else is
real drift and the command prints per-line descriptions like
step 2: status changed 'ok' -> 'failed'.
Same signal from Python:
Excluded from the distilled shape on purpose: wall-clock timestamps,
per-step duration_ms, frames filenames, resolved URLs (including
query-string tokens), cursor_xy. If you need those in your gate too,
diff the raw sidecar with your own tooling — the assertion set is the
narrow "did the UI still behave" contract, not the wire-level snapshot.
clickcast diffassertions is structural — it never looks at a pixel. clickcast diff
(and Reel.visual_diff()) is the pixel-level companion: it pairs up two
sidecars' steps and pixel-diffs their frames, reporting a percent-changed
and a list of changed bounding regions per step, plus region-highlighted
diff images. Reach for assertions to catch "the flow broke" (wrong step
count, a step that started failing); reach for diff to catch "the flow
ran fine but the button moved / the color changed / the layout shifted."
The two compose — run both in CI for full coverage, or diff alone if
pixel drift is your only concern.
Exits non-zero when any step's changed pixel percentage exceeds
--fail-above, or when a step couldn't be paired with its baseline
counterpart at all (mismatched step counts fall back to label matching;
anything still unmatched is flagged rather than silently skipped).
--out collects the region-highlighted diff images plus a summary.json
for CI artifacts; --threshold tunes the per-pixel noise floor. Clickcast's
own overlays (progress bar, action label, actions panel, cursor + ripple)
are excluded from the diff by default — otherwise every run would flag its
own chrome as a regression — pass --no-exclude-overlays for a strict
raw-pixel diff. Same signal from Python via reel.visual_diff(run_path, baseline_path).
Precedence (highest → lowest):
meta: blockCLICKCAST_* environment variables./clickcast.tomlclickcast config path)Every Config field can be set at any of these layers: engine, viewport, device, headful, slowmo, lang, dark, proxy, fps, dwell, format, quality, loop.
Project TOML — flat or [defaults]-wrapped both work:
Env vars:
| Format | Best for | Notes |
|---|---|---|
gif | READMEs, chat, quick shares | Widest compatibility; larger files. |
mp4 | Docs sites, social, long tours | Smallest for length; uses imageio-ffmpeg's bundled binary. |
webp | Web embedding | Great size/quality; animated. |
frames | Custom pipelines | Numbered PNGs + a frames.json manifest. |
--quality 1..30 trades size for fidelity (lower = better). --loop 0 loops forever; --loop 1 plays once.
click, type, scroll, …) with normalised timings and cursor tracking.The annotator (clickcast.annotate.Annotator — click ripples, cursor trail, caption bar, progress bar) ships as a library API in v0.1. Automatic wiring into auto / run outputs is planned for v0.2 (see Roadmap).
--initial-wait, or add wait: networkidle (or a specific selector) to the first step.ffmpeg not found — imageio-ffmpeg bundles a static binary; falls back if missing. Choose gif / webp if you'd rather skip MP4 entirely.clickcast elements <url> shows what's actually clickable. Or mark the step optional: true.CLICKCAST_PROXY, or proxy in the scenario meta: block.clickcast install. On Linux CI add --with-deps.src/clickcast/feedback/schema/v1.json; a future v2 (see #29) will add a graph block without breaking v1 consumers.Before opening a PR:
Cutting a release is documented in RELEASING.md.
v0.1 (this release): Session · Actions · Recorder · Encoder · Discovery · YAML scenarios · CLI · Python API · Sidecar (schema v1) · Config precedence · Fixture test site · Docs.
v0.2 (planned — tracked in #29):
clickcast explore <url> treats the app as a state graph: discover → click → discover the new state → recurse. Bounded, deterministic, with visited-state dedup.graph block (nodes = distinct page states, edges = (from, to, action, transition_kind)).auto / run outputs.vercel-labs/webreel — a TypeScript tool for authoring polished demo videos. Easy to confuse with this project given the similar name and overlapping output, so: clickcast is a Python tool aimed primarily at AI agents that need a visual modality onto a live web UI, and secondarily at humans who want reproducible demo reels. If you want hand-authored marketing videos, webreel is the better fit.
MIT © 2026 Alex Kay. See LICENSE.