claude-phantom

your app crashed. phantom is on it.

Wrap any command. If it dies, a headless Claude Code session diagnoses the bug, writes a failing test, patches it on a separate branch, verifies the fix independently, and leaves a post-mortem. Your branch is never touched.

$npm install -g claude-phantom
$phantom npm run dev
v0.7.0 555 tests 0 dependencies node ≥18 MIT
zsh — phantom npm run dev
$
TypeError: Cannot read properties of undefined (reading 'email')
at formatOrderLine (src/report.js:9:31)
 
👻 npm run dev crashed (exit 1) — phantom is taking over
phantom › working on branch phantom/fix-typeerror-…-k3f9a

the ninety seconds

Ten steps, and you can
stop it at any of them.

Phantom runs your command untouched and stays out of the way. Your exit code passes through unchanged, so it is safe inside && chains and in CI. Scroll to watch a recovery happen.

  1. 01Your command runs untouched. stdout, stderr and stdin stream through byte-for-byte; a ring buffer keeps the tail.
  2. 02It exits non-zero, or dies from a signal that was not your Ctrl+C.
  3. 03Phantom captures the crash — redacted output, stack trace, the files it names, git state.
  4. 04A phantom/fix-* branch is cut from HEAD. Nothing before this point has changed your tree.
  5. 05A headless Claude Code session starts, with an explicit allowlist, deny rules, and the guard hook watching every tool call.
  6. 06It writes a failing regression test first, then patches until that test passes.
  7. 07Phantom runs your test command itself, outside the session. Claude's claim is never trusted.
  8. 08The crashed command is re-run. Still crashing means unfixed, whatever the suite says.
  9. 09A post-mortem is written, the never-touch audit runs, and you are put back on your own branch.
  10. 10Your original exit code is returned. A fixed crash is still exit 1.
zsh — phantom npm run dev
$ phantom npm run dev
ready — listening on http://localhost:3000
TypeError: Cannot read properties of undefined (reading 'email')
at formatOrderLine (src/report.js:9:31)
👻 npm run dev crashed (exit 1) — phantom is taking over
phantom › captured → .phantom/crashes/2f9c.json
phantom › branch phantom/fix-typeerror-…-k3f9a cut from HEAD
phantom › ⠹ 1/3 · session started · allowlist + guard hook active
phantom › ⠧ 1/3 · wrote test/report.regression.test.js · 0m 41s / 15m
phantom › re-running your test command — outside the session
phantom › tests pass
phantom › npm run dev now exits 0 — the crash is gone
phantom › post-mortem → .phantom/reports/2f9c.md
phantom › never-touch audit clean · you are back on main
👻 phantom ✅ fixed · 1m 48s · 34.1k tokens
review git diff main..phantom/fix-typeerror-…-k3f9a

The spinner lines are live: the guard hook sees every file the session touches, so it tells you what is happening rather than only that something is.

the part you actually care about

It never commits
to your branch.

Letting an agent loose in your repo is the objection, and it should be. Every rail below is a mechanism, not a promise — and each one is there because the alternative was tried and written down.

the rails6 mechanisms · enforced, not promised
never your branch

A phantom/fix-* branch is cut from HEAD before any edit, and you are checked back out when it finishes — success or failure. The fix exists only as a branch to diff, merge, or delete.

never your secrets

.env, *.pem, *.key, secrets/** are enforced three ways: permission deny rules, a PreToolUse guard hook on every call, and a post-session audit that hard-reverts the branch on any hit.

never the network

No WebFetch, no curl, no git push. There is no push code path to configure.

never a green lie

Phantom re-runs your tests itself, outside the session, and then re-runs the command that crashed. A session that claims success without changing anything is reported unfixed.

never unbounded

maxIterations, maxMinutes, and a real spend ceiling — maxTokens or maxCostUsd — checked before each additional attempt.

Ctrl+C is a kill switch

Kills the process tree, rescues untracked work into a stash, resets the fix branch, checks your branch back out, exits 130. SIGTERM and SIGHUP do the same.

not a sandbox

The session runs node — it has to, to run your tests — so a one-liner can read what your user can read. The guard is lexical: an audit found four ways past it in one afternoon. All four are fixed and pinned by tests, and the honest lesson is that a lexical guard is a speed bump. Branch isolation and the post-session audit are the real backstops. Need hard isolation? Run it in a container.

redaction is pattern-based

Output, and the command line phantom displays, are scrubbed before the model, the report, the notification or the webhook ever see them — KEY=value, Authorization headers, URL credentials, PEM blocks, well-known token shapes. Unusual formats get through. It is a safety net, not a guarantee.

the surface

Five commands is
the whole thing.

Wrap anything. Flags go before the command; everything after it passes through verbatim, so phantom never has an opinion about your arguments.

$ phantom --max-minutes 10 npm run dev
ready — listening on http://localhost:3000
dry run

--dry-run diagnoses and proposes a diff with no branch and no edits. The CI-safe mode — and phantom now checks afterwards that the session really did change nothing.

every flag negates

--commit, --prompt, --no-notify, --verify. A setting in your config file can be turned off for a single run without editing it.

configured anywhere

.phantomrc, a package.json field, or fourteen PHANTOM_* variables. Precedence: flags > environment > file > defaults.

what it is not

The honest part.

Nothing here is a caveat buried in a footnote. If one of these is a dealbreaker, better you find out now than after a recovery.

git required

Outside a repo, phantom passes the command through and says why it is not recovering. There is no way to undo a bad session without git.

non-deterministic

Claude may fail. You get unfixed or timeout, your branch untouched, and a report saying what was tried.

best on node/js

Patching works anywhere Claude Code can edit, but verification needs a test command and the crash heuristics are tuned for Node traces first.

it costs real money

Every recovery is a genuine session on your plan or API key. Set maxCostUsd if that matters — the dollar figure phantom shows is an estimate from published rates, not your bill.

daemons keep it waiting

A child that never closes stdout keeps phantom waiting. Wrap foreground processes.

why trust it

It was audited harder
than it advertises.

Phantom is a tool that edits your code while you are not watching. That only deserves trust if the thing itself is held to a standard, so it is: every fix ships with a regression test, and every test is broken on purpose to prove it fails before it is allowed to pass.

532tests, run on 3 platforms × Node 18/20/22/24
53findings from a full audit of the tool itself
0runtime dependencies, by rule
15CI jobs green before anything ships

The release notes name what was broken rather than burying it — including the release where following phantom's own printed recovery instructions could destroy your uncommitted work. That bug is fixed, pinned by a test that executes the printed advice verbatim, and written up in full in the changelog.

questions

Reasonable doubts.

Does it ever push?

No. It is a denied tool, there is no network tool, there is no push code path, and there is no flag. Not configurable.

Can it touch my .env?

No — enforced as permission deny rules, by the guard hook on every call, and by a post-session audit that hard-reverts the branch on any hit. Watch your own log output though: the redactor is pattern-based.

I'm mid-change.

Phantom refuses on a dirty tree. --allow-dirty stashes a snapshot first and restores it after — including on Ctrl+C — and the command it prints to restore it names the stash by full sha, so a stash pushed meanwhile by another shell cannot be popped by mistake.

How much does it cost?

One to three headless turns plus test runs — roughly a short interactive debugging session. maxIterations and maxMinutes bound how often it asks and how long it waits; for an actual ceiling on spend, set maxTokens or maxCostUsd.

Can I use it in CI?

Use --dry-run: diagnosis and a proposed diff land in .phantom/reports/ with no branch and no edits. Upload .phantom/ as an artifact. Full mode works, but the branch dies with the runner since nothing is pushed.

From inside Claude Code?

Yes, and it does the right thing: a tool call times out long before a recovery finishes, so phantom captures the crash and hands it to /phantom:recover instead of starting a session that would be killed halfway. The recovery session also never inherits the parent session's environment.

What is the overhead when nothing crashes?

A child-process spawn with piped stdio and a bounded buffer. A 50 MB log flood is in the test suite.

Wrap one command.
Forget about it.

Then run phantom doctor once, before your first crash.

$npm install -g claude-phantom
$phantom doctor

Also installable as a Claude Code plugin — /plugin marketplace add waazy-w/claude-phantom