ADD is one skill file. Before any code, the agent writes the rules, its guesses and the checks that must fail — and seals them in a git commit. Then it builds to green and proves it. Here is what that measurably changes, and what it costs.
Nothing that matters lives in the chat. A fresh session reads the task file and git log and loses nothing; a reviewer diffs the seal instead of trusting a summary.
We seeded 24 bugs into each finished app and counted how many the agent's own tests caught. Every ADD check has to name the plausible wrong build it fails — and it shows.
Mutation score, 0–1. In round 5 on amb1 only one vanilla run wrote tests; that run scored higher.
One dot per run. marks a run that fell through the floor. Left alone, the agent sometimes ships no tests, stops to ask and delivers nothing, or crashes on a malformed request. With ADD it wrote its guesses down and kept going.
A held-out suite the agent never sees: null and number bodies, wrong types, timezone offsets. ADD's checks now send inputs the way a real caller sends them.
Held-out cases passed, mean of 3 runs (round 5). Part of this is taught: the skill names malformed bodies, and the suite tests them.
Every task ends with a verdict and the exact commands behind it. We reran each suite ourselves.
An earlier measurement (ADD 2.0.0, n = 1 per arm): over six evolving milestones, one long conversation let requirement coverage decay and never recover. A fresh session per milestone, resuming from files, held the line.
ADD is not free. It spends turns writing the contract before the code. On these workloads both arms pass the correctness oracle, so the gains above are in tests, robustness and trust — not in the pass rate.
We added four rules and measured each one. The rule that named something the agent must write — a check — changed its behaviour. Advice about how to think mostly didn't.
npx @pilotspace/add init
pip install pilotspace-add && pilotspace-add init
/plugin marketplace add pilotspace/ADD
| measure | vanilla | ADD 4.0 |
|---|
Read before quoting. Same model (claude-sonnet-5 --effort medium), n = 3 per arm per workload — direction, not proof. "Vanilla" is Claude Code carrying the operator's own ~/.claude config, which already asks for red/green TDD, so it is not bare Claude Code. Some runs were rerun after an account change; one run's cost is an estimate. Sources: the results page and the three pilot reports it links. Read the book.