Whippletree

Docs / Harnesses

Harnesses

Four targets today. They expose different events, not all of those events can block, and the differences decide what your contract can promise.

harnessbackendinstalltested against
claude-codehooks-jsonits own plugin marketplace2.1.0 and up
codexhooks-jsonits own plugin marketplace0.144.0 and up
copilothooks-jsonits own plugin marketplace1.0.80 and up
opencodets-pluginWhippletree writes the shim1.18.10 and up

What each event maps to

Where a harness has no equivalent, a requirement bound to that event lands ABSENT there, or REFUSES outright if you marked it hard-required. A blocking-gate or lifecycle-signal carrying a fallbackSkill is the one exception, and lands T3.

eventclaude-codecodexcopilotopencode
session-startSessionStartSessionStartSessionStartsession.created
session-endSessionEndSessionEndSessionEndnone
turn-endStop, blockingStop, blockingStop, not via exit codenone
tool-prePreToolUse, blockingPreToolUse, blockingPreToolUse, blockingtool.execute.before, blocking
tool-postPostToolUsePostToolUsePostToolUsetool.execute.after

Four primitives are missing from that table. No target maps subagent-start, subagent-stop, compact-pre or compact-post, so a requirement bound to one of them lands ABSENT on every target today, or REFUSES if you marked it hard-required.

What each tool class maps to

The file-read, file-write and shell-exec aliases resolve to these tools.

classclaude-codecodexcopilotopencode
readReadno such toolReadread
writeWriteapply_patchWritewrite
shellBashBashBashbash

The table hides a gap. file-write binds the write tool alone, so a change to an existing file misses it on claude-code and copilot (which route that through Edit) and on opencode (edit). On codex the alias binds apply_patch, and codex has no probe note of its own, so whether an edit to an existing file reaches it there is unverified.

claude-code

Every event in the table above maps to a native hook and every tool class names a real tool, so nothing degrades here. The plugin carries the skills.

codex

Its events match claude-code's, but it has no file-read tool. The target definition keeps file-read by declaring a degradation to T2:

file-read  want ≥T4  got T2  SATISFY   matcher Bash|Edit|Write|apply_patch;
                                       misses reads in pipelines/heredocs/scripts;
                                       may double-count

The target definition carries that lossage, so preflight prints it before you install instead of after an install has already gone wrong. If T2 is not good enough, raise minTier to T1: a hard-required requirement then REFUSES on codex, a soft one lands DEGRADE.

copilot

Copilot accepts the file Whippletree already emits, unchanged, so it needed no new backend. Its own documented entry format is flatter, but the nested Claude shape works, and that is what ships.

Whippletree cannot drive Copilot's turn-end gate yet, and the limit is in the dispatcher. A handler exiting 2 on Stop is ignored and the session ends; the same handler exiting 2 on PreToolUse denies the call outright, which is the control that makes this a measurement. Copilot does honour a {"decision":"block"} written to stdout on Stop: the hook re-fires and the agent acts on the reason.

A hard-required blocking-gate on turn-end REFUSES on copilot today unless it declares a fallbackSkill and accepts T3. The reason differs from opencode's. opencode has no blocking stop event; Copilot has one Whippletree cannot yet drive, because the dispatcher signals a block by exiting 2 and nothing else. Reporting that as REFUSE rather than as a working gate is the honest answer, and closing that gap would make this T1. Tracked as issue 19.

One trap: hooks are not auto-discovered, and Copilot reads .plugin/plugin.json before the root plugin.json that every bundle already ships. Had it been the other way round, a bundle would have installed cleanly and registered nothing at all.

opencode

opencode has no blocking stop event. A thrown error fails a single tool call and the agent loop carries on, so there is nothing to bind turn-end to.

A hard-required blocking-gate on turn-end REFUSES on opencode unless it declares a fallbackSkill and accepts T3. That is the correct answer, not a limitation to work around: a gate that degrades to advice without saying so is worse than one that will not install, because you would believe it was running.

If advice is acceptable, pair the gate with a skill through fallbackSkill: the step compiles into instructions, and a minTier of T3 turns that REFUSE into a SATISFY. opencode also uses a different backend: it has no plugin manifest to extend, so install writes a TypeScript shim into .opencode/plugin/ and copies skills into .opencode/skills rather than shipping them in the bundle.

Windows

Windows is a property of the machine, not another target: no target definition changes there, and what changes is the handler. A handler written as a shell script does not work, because Windows has no shebang support. A requirement can declare handlerWindows alongside handler, and the dispatcher picks per platform at run time. Where a requirement declares none, the dispatcher skips it on Windows and says so on stderr.

A handlerWindows has to be something Windows launches from a bare path, since the dispatcher execs it with no interpreter: .exe, .com, .bat or .cmd. .ps1 is not one of them: it is absent from the default PATHEXT, so wrap it in a .cmd. contract.Validate rejects anything else, so build, preflight and install all refuse it instead of leaving it to fail open at dispatch.

Preflight does not model the platform. It reports what a harness can carry, and a harness is not a machine. It still reports SATISFY for a requirement with no handlerWindows, which dispatch then skips on Windows.

No harness has been installed on Windows. The rules above come from spawning the dispatcher and handlers directly on a Windows runner, so they say what the loader launches, not what a given harness does with it. Running a bundle there is also unfinished: init scaffolds only shell handlers, and the T3 instruction fallback is written as a POSIX shell snippet. Tracked as issue 1.

How these claims are checked

Every row above comes from a target definition in the repository, and each definition declares the harness versions it is tested against. Most rows were measured against a harness running in a sandbox. Where a probe could not provoke an event, or where a tool name came out of the harness's source instead of a run, the notes say so, and they live beside the code. Codex has no probe note of its own: its rows come from reading the harness's source and from the nightly job. That job reinstalls claude-code, codex and opencode and reruns the end-to-end suite, which watches installation, session-start, preflight verdicts and the dispatcher's own output, so a harness moving under those shows up as a failure instead of as a wrong answer in your terminal. Copilot is not in the job yet: the events its script asserts on only fire once an agent takes a turn, so it needs a real login.

Next