ADD 4.0 · measured, rounds 3–5 on Sonnet 5 · round 6 on Sonnet 5.5: the two flows, side by side
Claude Code with and without the ADD skill · same pinned model · n = 3 per arm per workload

Claude Code writes the code.
ADD makes it something you can trust.

ADD is one skill file. Before any code, the agent writes the rules, its guesses and the checks that must fail — and seals them in a git commit. Then it builds to green and proves it. Here is what that measurably changes, and what it costs.

0.79
share of 24 seeded bugs its own tests caught (amb1)
vanilla: 0.53
27 of 27
runs whose claimed test count matched a fresh rerun
vanilla makes no claim to check
0 of 13
runs that stalled on a contradictory spec
vanilla: 2 of 7 shipped nothing
2.2–2.9×
the dollars per run — the price of the above
correctness oracle: both arms pass

The loop — every beat leaves something on disk

Nothing that matters lives in the chat. A fresh session reads the task file and git log and loses nothing; a reviewer diffs the seal instead of trusting a summary.

Directionrules · guesses · red checks Buildto green · seal untouched Verifyfresh run · probes · verdict
1 · freeze(booking-api): rules, assumptions, failing checks — sealed
2 · feat(booking): build to green — sealed files untouched
3 · verify(booking-api): PASS — commands, exit codes, counts in the task file

Tests that catch real bugs

We seeded 24 bugs into each finished app and counted how many the agent's own tests caught. Every ADD check has to name the plausible wrong build it fails — and it shows.

Claude Code + ADD 4.0Claude Code, vanilla

Mutation score, 0–1. In round 5 on amb1 only one vanilla run wrote tests; that run scored higher.

A floor that always holds

One dot per run. marks a run that fell through the floor. Left alone, the agent sometimes ships no tests, stops to ask and delivers nothing, or crashes on a malformed request. With ADD it wrote its guesses down and kept going.

Edge cases it never saw

A held-out suite the agent never sees: null and number bodies, wrong types, timezone offsets. ADD's checks now send inputs the way a real caller sends them.

ADD 4.0vanilla

Held-out cases passed, mean of 3 runs (round 5). Part of this is taught: the skill names malformed bodies, and the suite tests them.

Evidence you can re-run

Every task ends with a verdict and the exact commands behind it. We reran each suite ourselves.

27 of 27
claimed test counts matched a fresh rerun (15 + 6 + 6 runs across three rounds)

Why the record lives on disk — context rot

An earlier measurement (ADD 2.0.0, n = 1 per arm): over six evolving milestones, one long conversation let requirement coverage decay and never recover. A fresh session per milestone, resuming from files, held the line.

fresh session, resumed from diskone long conversation

The honest cost

ADD is not free. It spends turns writing the contract before the code. On these workloads both arms pass the correctness oracle, so the gains above are in tests, robustness and trust — not in the pass rate.

ADD 4.0vanilla

Where ADD's tokens go

What changed the model — and what didn't

We added four rules and measured each one. The rule that named something the agent must write — a check — changed its behaviour. Advice about how to think mostly didn't.

Try it on your Claude Code

npx @pilotspace/add init pip install pilotspace-add && pilotspace-add init /plugin marketplace add pilotspace/ADD
Every number on this page, as a table
measurevanillaADD 4.0

Read before quoting. Same model (claude-sonnet-5 --effort medium), n = 3 per arm per workload — direction, not proof. "Vanilla" is Claude Code carrying the operator's own ~/.claude config, which already asks for red/green TDD, so it is not bare Claude Code. Some runs were rerun after an account change; one run's cost is an estimate. Sources: the results page and the three pilot reports it links. Read the book.