20 · What changed in 4.0 — migrating from 3.x¶
← 17 Components — monorepo and multi-repo · Contents · Next: Appendix C Glossary →
In one paragraph¶
ADD 4.0 removes the engine. There is no .add/tooling/, no cli.py, no add.py, no add <verb> commands, no refusal codes and no agent roster. The method is one skill file — SKILL.md with four references: format.md (file shapes), explore.md (the research lane), evidence.md (the closed loop from intent to production) and personas.md (routing lenses) — that the agent follows with the tools every project already has: files, git, and the project's own test command. There is also no approval step: the agent routes, seals, builds and verifies without stopping, and the human reviews afterwards from a report that lists every assumption it took.
Why¶
The engine was built to make the method's promises mechanical: a freeze could not be faked, a verdict needed a receipt, a stale green was refused. It did that. It also cost more than it protected.
- Cost against vanilla prompting. On the amb1 ambiguity track, vanilla prompting cost $0.61 per run with an oracle score of 1.00; ADD 3.x cost $2.13–2.52 per run with oracle scores of 0.75–1.00 (amb1 findings). The intervention that run measured changed what ADD wrote down, not what it built.
- Cost against spec-kit. Over five evolving milestones, ADD 3.x cost $34.37 against spec-kit's $18.66 — 1.8× in dollars — at a minimum per-milestone fidelity of 0.97 against 0.95. On the two larger milestones ADD's cost rose far more steeply, driven by per-task ceremony over a growing codebase (benchmark, WM4/WM5 extension).
- Where the tokens went. Cost tracked turns × context per turn, and a large share of the agent's turns were spent driving the engine — status, stamps, receipts, gates — rather than doing the work.
What the engine enforced turned out to be expressible without it:
| 3.x mechanism | 4.0 equivalent |
|---|---|
add freeze stamping a digest of the direction |
a freeze(<slug>) git commit of the task file and its check files |
| the tamper tripwire on sealed files | git diff <freeze> HEAD -- <sealed files> must print nothing |
add refreeze / change requests |
a refreeze(<slug>): <why> commit, reason under ## LOG |
add run writing a Run receipt |
real command output — command, exit code, counts — written into ## EVIDENCE |
add gate PASS \| RISK-ACCEPTED \| HARD-STOP |
the same three verdicts, written into ## EVIDENCE, committed as verify(<slug>): <verdict> |
add refute stamps and refute tiers |
1–3 executable probes from the sealed rules, written after the build under the counter-lens persona — by a fresh subagent for security work, by a cold reread otherwise |
add interview and the human freeze |
## ASSUMPTIONS and derived: rules, reviewed afterwards in the session report; for floor tasks a second reader under the counter-lens challenges them before the seal — a fresh subagent for security work, a cold reread for the rest |
| a Must names its source | (from: request \| <file> \| <spec> \| derived: <why>) on each rule; each check names its falsifier: |
| the regression floor PLAN line | regression: in PLAN — the full suite, affected: … — why, or none — why |
a refreeze that moves gives: marks consumers stale |
Verify step 3: git grep the surface's users and run their tests; a broken consumer is not a PASS |
| the quick-lane tripwire | Quick work that turns out to touch the floor "is a Task now" — stop and write the task file |
add release <tag> binding a tag to receipts |
tag only a commit whose tasks since the last tag each end in a verify( commit; observes: names what to watch after release |
| an escape drains only with why-missed + prevention; closed history superseded | a successor task with fixes: <slug>@<verify sha>, a reproducing check, why the checks missed it, and a bound prevention |
add status |
read .add/PROJECT.md and the open task files; git log |
add new |
write the task or milestone file from format.md |
add learn / add deltas / add fold |
a line in the spec's ## Deltas; promotion to ## Decisions that bind |
add milestone-done goal gate |
tick each EXIT box with its evidence; status: done when all are ticked |
add wave / add join |
one git worktree per independent task, disjoint scope:, merged one at a time |
add advise / persona records |
route a lead persona by the task's risks: against each persona's covers-risks:; its counter-lens: reads at verify; a lens: line in EVIDENCE records what it caught |
add doctor --sync, graph.json, compiled index.md, log.md |
nothing — nothing is compiled; the files are the state |
add-worker / add-advisor agents |
removed; the model plans itself and spawns at most one foreground subagent per beat, for security work |
What did not change: Direction → Build → Verify; rules, assumptions and checks before any production code; red for the right reason; never weaken a sealed check; the three verdicts; security is always a HARD-STOP; invariants: bind every change; the lanes and the floor.
Measured on 4.0¶
The engine's removal was argued from cost. What the skill buys was then measured against vanilla Claude Code — same pinned model, n = 3 per arm per workload, the operator's ~/.claude config loaded by both arms (results):
- Stronger tests. The share of 24 seeded bugs the agent's own tests catch: 0.68 against 0.51 on the greenfield workload, 0.79 against 0.53 on the ambiguous one. This is the falsifier on every check at work.
- A floor that always holds. ADD shipped tests in 12 of 12 runs; vanilla built code with no tests in 2 of 5. On a contradictory spec vanilla stopped to ask and shipped nothing in 2 of 7 runs; ADD recorded its choice as an assumption and delivered in 13 of 13.
- Honest evidence. The test count ADD claimed matched a fresh rerun in 27 of 27 runs.
- The cost. 2.2–2.9× the dollars. The correctness oracle passed in both arms, so these workloads cannot show a correctness gain either way.
- On Sonnet 5.5 (round 6). The same skill on a newer model cost 1.7–2.1× vanilla's dollars and 4.0–5.1× its minutes, and its Direction beat fell from 4.1 min to 1.1–1.7 min. Vanilla now ties on every held-out edge case, and mutation scores sit within noise. What still separates the arms is how the silences get read: 5.7 of 7 planted ambiguities handled right against 4.3, and "who may cancel?" read as owner-only in 3 of 3 runs against 0 of 3. Watch the two flows side by side.
One lesson shaped the skill itself: a rule changes behaviour when it names something the model must write — a check, a covers: line. Checks that send inputs the way a real caller does ended malformed-body crashes (4 of 6 runs → 0 of 6); advice about how to guess, order a report or batch tool calls barely moved.
Upgrading a project¶
- Update the install with the same command you installed with — e.g.
npx @pilotspace/add@latest update, orpipx run pilotspace-add update. The installer replaces the skill, removes.add/tooling/and the retiredadd-worker/add-advisoragent files, and seeds any missing starter personas without overwriting yours. - Keep your bundle. A 3.x
.add/reads as-is. Its extra frontmatter (verified:,generated:, digests),graph.json,index.md,log.md,runs/andtasks/<slug>.d/are history: leave them, do not maintain them. New work uses the shapes in 12 · The bundle. If you prefer a clean tree, move the 3.x bundle aside — this repository keeps its own underarchive/add-3x-bundle/. - Make sure
PROJECT.mdhasgoal:,invariants:andtest_cmd:. The agent reads it first every session. - Open tasks carry on under the 4.0 loop. A task frozen under 3.x has no
freeze(<slug>)commit; when resuming it, commit its task file and check files asfreeze(<slug>)so Verify has a seal to diff against. - Update your agent pointers. Re-running the installer rewrites the managed ADD block in
CLAUDE.md/AGENTS.md; text outside the markers is untouched. - Retire scripts and CI steps that call the old CLI. A CI job that ran the engine's audit becomes an ordinary test run.
What you give up¶
Mechanical refusal. In 3.x the engine refused to record a PASS over a stale receipt; in 4.0 a dishonest agent could write one. What stands in its place is that every claim is checkable by anyone with git: the seal is a diff, the evidence names the exact commands at an exact commit, and every change of intent is a visible commit. The human's review moved to the end, where the evidence is — and the report is written so that disagreeing is easy.