The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Isitdone listing page.
Don't let your coding agent say "done" until the tests actually pass.
Rendered from a real run (npm run demo): the agent lines are narration; the hook and CLI output are captured as-is, with only the temporary path shortened.
isitdone is a zero-LLM, zero-dependency Stop hook and CLI for Claude Code, Codex CLI, Cursor, Gemini CLI, GitHub Copilot CLI, Qwen Code, Goose, Factory Droid, Devin, Augment, OpenCode and Junie CLI. When the agent tries to end its turn claiming the work is complete, isitdone runs the repository's real test, typecheck and lint commands on the exact working tree, scans the diff for weakened tests, and refuses the stop until they pass. It also tells the agent mid-turn when an edit just weakened a test, runs the same verification on pull requests as a GitHub Action, offers it as an MCP server to agents that have no stop hook, and leaves a git-bound receipt you can paste into a PR.
No API keys. No network. No telemetry. Just exit codes.
Claude Code ends its turn with:
Done. All 4 tests pass and the auth refactor is complete.
The Stop hook runs isitdone before that turn is allowed to end. The agent receives this instead:
The agent keeps working. When it genuinely finishes:
Edit one more file and the receipt goes STALE until the checks run again. A hand-edited receipt reads NONE. (isitdone on npm is a short alias of @aivolution/isitdone; both commands are the same program.)
Measure it on your own machine. isitdone history reads the transcripts already on disk: Claude Code (~/.claude/projects), Codex CLI (~/.codex/sessions: legacy, paginated and pre-0.40 rollouts, honouring /undo rollbacks), Gemini CLI (~/.gemini/tmp/<project>/chats, both the .json and the 0.39+ .jsonl layouts), Qwen Code (~/.qwen/projects/<cwd>/chats) and Cursor (the IDE's state.vscdb bubble store, opened read-only in place, which needs Node 22.13+ or 24 for node:sqlite; on Node 20 the agent-transcripts JSONL is read instead, which carries no exit codes, and the report says so). It finds every turn where the agent edited files and then claimed completion, and checks whether a test command actually passed after the last edit. Nothing leaves your machine; only counts are printed.
That is the author's real result, printed by isitdone 0.8.2 on 2026-09-27 over the Claude Code sessions that were on disk before the gate went in: claims dated 2026-06-20 to 2026-09-07, which --until 2026-09-07 reproduces. One private project is left out with --exclude, as it was in every published run; the only other edits are the home directory shortened to ~ and the per-project lines at the end dropped. The first published figure came from isitdone 0.2.0 on 2026-09-07: 69% of 516 claims (31% verified, 37% stale, 31% never ran, and a failing-run row printed as 0%; under 0.2.0's plain rounding that means at most 2 claims). Later versions detect claims and turns more strictly, over the same transcripts. Since 0.8.1 history prints the count beside every share (since 0.8.2 the headline too) and never rounds a row that holds anything to 0%, so a 0% row now means none. Post yours.
Once the gate is installed, that measure alone undercounts: the verification happens inside the Stop hook, and a transcript does not show the hook's run as a test the agent ran. So the hook keeps a local log of its own decisions (.isitdone/decisions.jsonl: one line per stop with the outcome, whether the final message claimed completion, and check counts; no message text, no paths, the session id only hashed), and history adds what the gate did in the projects it scanned. The layout, with illustrative numbers:
The longer story of why the gate sits where it does is in docs/why.md (also on DEV, where comments are open).
One command per host. Run it inside the repository.
| Host | Command | Where it writes |
|---|---|---|
| Claude Code | npx isitdone init | .claude/settings.json (Stop) |
| Codex CLI | npx isitdone init --agent codex | .codex/hooks.json (Stop), then run /hooks in Codex and trust it |
| Cursor | npx isitdone init --agent cursor | .cursor/hooks.json (stop) |
| Gemini CLI | npx isitdone init --agent gemini | .gemini/settings.json (AfterAgent); or gemini extensions install https://github.com/raimondasl/isitdone-gemini |
| GitHub Copilot CLI | npx isitdone init --agent copilot | .github/hooks/isitdone.json (agentStop); restart Copilot |
| Qwen Code | npx isitdone init --agent qwen | .qwen/settings.json (Stop) |
| Goose | npx isitdone init --agent goose | .agents/plugins/isitdone/hooks/hooks.json (Stop) |
| Factory Droid | npx isitdone init --agent droid | .factory/hooks.json (Stop) |
| Devin | npx isitdone init --agent devin | .devin/hooks.v1.json (Stop); skipped when the Claude Code hook is present, since Devin loads that too |
| Augment (Auggie) | npx isitdone init --agent augment | .augment/settings.json (Stop) pointing at .augment/hooks/isitdone-hook.sh/.cmd, since Auggie runs script files |
| OpenCode | npx isitdone init --agent opencode | .opencode/plugins/isitdone.js (a plugin: OpenCode has no blocking hook, so failed checks come back as a visible [isitdone] follow-up message in the same session) |
| Junie CLI (early access) | npx isitdone init --agent junie --user | ~/.junie/config.json (Stop) |
| Everything | npx isitdone init --agent all | all of the above that apply to the repo |
init detects the checks, writes the Stop hook and (for Claude Code, Codex, Gemini CLI, Qwen Code, Devin and OpenCode) a warn-only post-edit hook, adds .isitdone/ to .gitignore, and runs doctor, which pipes a synthetic "all tests pass" stop event through the hook and proves it blocks, and runs the detected checks once on the current tree so that a wrong guess (say mypy . where CI runs mypy src/pkg) or an already-red tree shows up now, at install time, instead of as NOT DONE on every stop:
Claude Code users can also install it as a plugin: /plugin marketplace add raimondasl/isitdone then /plugin install isitdone@isitdone. Gemini CLI users can install it as an extension: gemini extensions install https://github.com/raimondasl/isitdone-gemini (isitdone-gemini). A paste-to-agent version of these instructions is in docs/install.md. How each of these agents' end-of-turn hooks works (config file, payload, how to block, loop flags) is written up vendor-neutrally in docs/agent-hooks.md, with the same table as data in docs/agent-hooks.json.
Add --user to install into your user-level settings instead of the project. npx isitdone uninstall removes it. Teach the agent to run it itself with npx skills add raimondasl/isitdone (the SKILL.md is at the repo root).
The code lives in the @aivolution/isitdone package; isitdone on npm is a short alias with the same command, and installed hooks always call the canonical package. npm i -D @aivolution/isitdone makes the hook resolve locally, with no registry lookup and offline.
The hook runs through npx, which keeps its own install cache and does not refresh an unpinned package by itself. To move the hook to the latest release:
doctor reports the version the hook actually runs and mentions when a newer release is available. Projects that installed @aivolution/isitdone as a dev dependency update it with npm update @aivolution/isitdone instead.
The Stop hook is the gate; the post-edit hook is the nudge. On Claude Code, Codex, Gemini CLI, Qwen Code, Devin and OpenCode, init also registers a hook that runs after every Edit/Write (Codex and Devin: apply_patch, Gemini and Qwen: write_file/replace, OpenCode: edit/write/apply_patch). It scans just that file against HEAD (so it sees every uncommitted change to the file, not only the lines this edit touched) and, when the tests got weaker, adds a short factual note next to the tool result:
It never blocks and it is fast (one file, no test run). Cursor has no channel an agent can see after an edit, so there the warning arrives with the Stop hook instead. Skip it with init --no-edit-hook; "integrity": "off" in the config turns it off as well.
The same verification on every pull request, from a clean checkout that never trusts a local receipt:
Listed on the GitHub Marketplace as isitdone verify. It runs the repo's checks, scans the diff against the PR base for weakened tests (strict by default: high/critical findings fail the job even when the checks pass), writes the receipt to the job summary, keeps one updated comment on the PR, and can upload the findings as SARIF (with: { sarif: 'true' }, needs security-events: write on push events). On pull requests from forks the default token is read-only, so the comment is skipped and the job summary carries the receipt; pass a token with write access (a PAT or a GitHub App token) or run on pull_request_target to comment there too. base accepts a branch (default: the PR base), a commit sha, a tag or a qualified ref. Inputs: version, base, strict, comment, sarif, token, args, command; outputs: status, report, json (the path of the --json-file result).
For agents and IDEs that have no stop hook a program can block (VS Code Copilot agent mode, Cline, Windsurf Cascade, Kiro, Zed, JetBrains AI Assistant, Claude Desktop, Amp, Crush, Roo Code), the same verification is available as a Model Context Protocol server:
This is the weaker of the two integrations, and it is meant to be. A hook is enforced by the host: the agent cannot end its turn until the checks pass, whether it wants to or not. An MCP tool runs only when the model decides to call it, and a model that skips the call claims "done" exactly as before. If your agent is in the Install table, use the hook (you can run both: they share the receipt). Where there is no hook, a tool the model is told to call still beats a sentence nobody checked.
It offers three tools, and its server instructions tell the model to call isitdone_verify before saying that work is complete, to paste the result, and never to weaken tests to make it pass:
| Tool | Arguments | What it returns |
|---|---|---|
isitdone_verify | cwd?, profile? (lite | full), claim?, base?, strict? | The npx isitdone run: done true/false, every check's status, the tail of the failing output, the test-integrity findings, the receipt state. NOT DONE is a normal result (isError stays false), so the model reads it and keeps working. |
isitdone_receipt | cwd? | PASS, FAIL, STALE or NONE for the current tree. Runs nothing. |
isitdone_detect | cwd? | The checks that would run, and where each was detected. Runs nothing. |
Every result is a plain-text report plus the same facts as structuredContent (with an outputSchema) for clients that read it. Without a cwd argument the server uses the client's first workspace root (roots), then the directory it was started in.
VS Code, in .vscode/mcp.json:
Cursor (.cursor/mcp.json), Claude Desktop (claude_desktop_config.json, under Settings, Developer), Cline (MCP Servers, Configure: cline_mcp_settings.json) and Windsurf (~/.codeium/windsurf/mcp_config.json) share one shape:
Claude Code:
Cursor and Claude Code have real hooks, so there the MCP server is a convenience (the agent can ask for the receipt mid-task), not the gate: keep npx isitdone init. On Windows, a client that cannot start npx directly takes "command": "cmd", "args": ["/c", "npx", "-y", "@aivolution/isitdone", "mcp"]. Claude Desktop has no workspace, so the model has to pass cwd. Since nothing forces the call, say it in the agent's rules file as well (.github/copilot-instructions.md, .clinerules, .windsurfrules, AGENTS.md): "Before you tell me work is done, call the isitdone_verify tool and paste its result."
Details that matter in practice: one verification runs at a time per repository, and a second identical call gets the result of the run in flight; when a client gives up on a slow suite (notifications/cancelled after its own tool timeout) the run still finishes and leaves its receipt, so the model's retry is answered at once instead of timing out again; progress notifications are sent per check for clients that ask for them; when the client closes stdin the running check is killed and the server exits. Checks run on pipes, so nothing a test prints can reach the protocol stream. The server is written against the wire format (newline-delimited JSON-RPC 2.0, no SDK, still zero dependencies) and speaks the handshake revisions 2024-11-05, 2025-03-26, 2025-06-18 and 2025-11-25 as well as the stateless 2026-07-28 revision (server/discover, per-request _meta, roots through input_required). server.json describes it for the official MCP Registry as io.github.raimondasl/isitdone; the release workflow publishes it there.
package.json scripts (test, typecheck, lint, build; npm, pnpm, yarn, bun, deno), pyproject.toml/pytest.ini/requirements.txt (pytest, ruff, flake8, mypy, pyright; uv/poetry/pipenv runners), go.mod (go vet, go test ./...), Cargo.toml (cargo check, cargo test), .NET solutions, Gradle/Maven, and Makefile targets. Anything can be overridden in .isitdone.json.claim-gated profile runs the fast lite checks (typecheck, lint) on every stop, and the full checks (tests, build) only when the agent's final message contains a completion claim: "tests pass", "done", "implemented", "verified", "ready for review", and so on. A question or a progress update does not trigger a two-minute test run. Hosts that do not pass the final message (Cursor, Copilot CLI, Factory Droid, Devin, Junie) get the full profile, cached per tree..git/objects. A PASS receipt for the same tree hash and the same check configuration is reused; nothing runs twice for nothing.CI=true, per-check timeouts, and the last 30 lines captured. If anything fails, the hook returns the host's block shape with a bounded, plain-text reason quoting the claim and the failing output. If everything passes, the receipt is written and the agent may stop..skip/.only/xfail, dropped assertions, matchers downgraded (toStrictEqual to toEqual, toThrow("msg") to toThrow(), assertEqual to assertTrue), widened tolerances, empty catch/except: pass, and neutered configuration (|| true, --passWithNoTests, continue-on-error, testPathIgnorePatterns, -DskipTests, ignoreFailures, cargo test -- --skip, removed CI test steps). JS/TS, Python, Go, Rust (#[ignore], #[should_panic] loosened, assert_eq! to is_ok(), inline #[cfg(test)] modules found by content), Java/Kotlin (JUnit 4/5, TestNG, AssertJ, Hamcrest; @Disabled, assumeTrue(false), assertEquals to assertNotNull; Kotlin is best effort for JUnit and kotlin.test) and C# (xUnit, NUnit, MSTest; Skip =, [Ignore], Assert.Equal to Assert.NotNull); the build configuration scanned includes Cargo/nextest, Maven/Gradle, csproj/runsettings/xunit.runner.json and the common CI files. Findings are reported with a before/after line (Tests 47 -> 44 Assertions 112 -> 104 Skipped 0 -> 1) and recorded in the receipt; with "integrity": "strict" (or --strict) high/critical findings block the stop even when the checks pass. Suppress a line with // isitdone: allow <reason>; suppressions are reported, never hidden, and --ci treats new ones as findings.stop_hook_active / loop_count (and counts attempts itself for hosts that send neither), counts its own attempts per session (default cap 3), and after the cap lets the agent stop with a visible warning. Malformed stdin, a broken config, or an internal error always allow the stop: isitdone must never brick the agent.The checks see the whole working tree. When two agent sessions work in the same directory, that tree holds the other session's unfinished files too, and a gate that ignores this blames whichever session stops first and tells it to "fix" work that is not its own. isitdone keeps the two apart, under one rule: a session on its own is gated as before. Anything softer needs proof that another session has work in progress here.
.isitdone/sessions/ at the repository top: paths and timestamps, local, never committed)./clear, a restart, yesterday's session) is not "another session", and a session that finished its turn with passing checks (or a one-shot helper that passed and exited) has nothing in progress: its files were good when it left them, so a failure in them now is a later change's doing, and the session that made that change is held to it in full.checkout, restore, reset, stash, clean).maxAttempts.What "names a file" means: the repo-relative path, an absolute path under the repository, or the shorter path a check prints when it runs from a sub-directory or a workspace, provided only one file in the repository ends that way; a bare Name.java:17 counts only for a distinctive name that is unique in the repository. Colour codes are stripped first, and on Windows and macOS the comparison ignores case.
Limits worth knowing. The records come from the post-edit hook, so ownership is known on Claude Code, Codex, Gemini CLI, Qwen Code, Devin and OpenCode (where a subagent's edits count as its parent's; re-run init --agent opencode once to refresh a plugin generated before 0.6.1); files changed through shell commands are nobody's. A session killed mid-turn never says goodbye: for up to an hour its unfinished files still soften the gate, to one block instead of three, for failures that name none of the remaining session's files. The reverse gap: a live session is invisible to a newcomer until one of its hooks fires after the newcomer's first. Attribution follows where a failure is reported, not what caused it: if this session's change breaks a file the other session is still editing, the single block and its question are the only safeguard, and the other session is held to the failure in its file. A file added while checks ran, by a session without a post-edit hook or through the shell, cannot be told from the checks' own output, so the receipt can still cover it unseen. A failing test usually names the test file, not the source file that was edited, so such a failure counts as "names none of its files" unless the session also edited the test. ../-relative paths in check output are not resolved. The release message is shown on Claude Code, Codex, Gemini CLI and Qwen Code; Devin and OpenCode have no channel for one and release silently after the one block. The run lock is best effort: when the wait runs out, or under heavy contention, runs can still overlap, as they always did before. "otherSessions": "ignore" switches off the instruction, the single block and the run lock (edits are still recorded; only the receipt binding reads them).
Sharing a directory this way is workable, not ideal: a whole-repo check cannot pass while the other session's half of the tree is broken. For long parallel work, give each session its own git worktree; each gets its own .isitdone/ and its own receipt.
For orchestrators and agent loops, npx isitdone --json is a done-predicate: done is true only when every full check passed on the current tree. Exit codes: 0 done, 1 not done, 3 usage or internal error.
Optional. .isitdone.json at the repo root, or an "isitdone" key in package.json:
kind: "lite" checks run on every stop; kind: "full" checks run when the agent claims completion. test and build default to full, typecheck, lint, check and vet default to lite.
isitdone guards honest mistakes, which is where nearly all "tests pass" fiction comes from.bench/ (npm run bench, which prints the version and commit it ran on). On 2026-09-24, at isitdone 0.7.0, it held 192 hand-labelled cases: no legitimate case flagged (0 of 89), 102 of 103 tampering cases caught, and every scored detector covered by at least three cases that name it alone; the run prints recall per detector as well as per case. It grows with every reported mistake.| Tool | What it does | Relation |
|---|---|---|
| oh-my-agent | Multi-agent framework that includes a stop-hook gate running typecheck/test/lint | isitdone is that gate as a standalone primitive for any host, with receipts and claim-gating |
| taskmaster | Blocks the stop until a completion token appears | Token-based; does not run the checks |
| gutcheck | Diff-scoped mutation probe with a Claude Code stop hook | Complementary second gate; isitdone verifies the suite passes, gutcheck verifies the suite means something |
| checkwash, testseal | Diff scanners for weakened tests | Diff-only; isitdone runs the checks (diff scanning is on the roadmap) |
| ProofRun | Tree-bound PASS/FAIL/STALE receipts with hand-written config | Similar receipt idea; isitdone auto-detects and hooks into the agents |
--related mode that runs only the tests touching the changed files for slow suites; Cline (a PreToolUse gate on attempt_completion) and Amp (agent.end plugin) adapters; a native OpenCode hook once session.stopping ships, and Windsurf/Cascade once its hooks can block; history for OpenCode's database; Kotest/Spek DSLs; a detector for expected values bent to match a regression.aider --auto-test --test-cmd "npx isitdone" feeds the same verdict back after every edit.The repository dogfoods itself: CI runs node dist/isitdone.js --json on every push, and pull requests run the GitHub Action. The library is importable too: import { verify, scanIntegrity } from '@aivolution/isitdone'.
MIT