ADD · the book · the measured tour
Round 6 · Claude Sonnet 5.5 · two real runs, condensed from their transcripts

Same request. Same model. Two flows.

Both runs got a booking-service spec that contradicts itself (a conflict is 202 waitlisted and 409 rejected) and says nothing on six other decisions. Watch where each one puts its decisions, and what that costs.

t = 0 s
🏃 vanilla
…
🛡️ + ADD
…
0 s50 s100 s150 s200 s
🏃 vanilla Claude Codevanilla-amb/rep1
waiting to start

In the repo afterwards

🛡️ Claude Code + ADDadd-4-amb/rep1
waiting to start

In the repo afterwards

Where did the decisions go?

Both runs built working code, and both caught the contradiction. The difference is where each decision lives once the chat is gone, and the guess nobody asked about.

🏃 vanilla
🛡️ + ADD
The 202-vs-409 contradiction
Caught it. Chose waitlist-by-default with a "waitlist": false opt-out, in its chat reply.
Caught it. Made the same choice, as ASSUMPTION A1 in the task file, listed first in the report as the costliest guess.
Who may cancel a booking?
Anyone. The reply never raises it.
The owner only: rule R:OWNER, pinned by sealed check C11.
Left in your repo
The code and 7 tests.
The code, 19 checks, and a task file with 17 rules, 8 assumptions and the evidence, as three commits: freeze, build, verify.
Planted ambiguities handled right
5 of 7
6 of 7

Over three runs each: what you get, what you pay

One pair of runs can mislead, so the board shows means over n = 3 runs per arm, on both workloads: wm1, a booking API built from scratch, and amb1, the contradictory spec above.

vanilla Claude CodeClaude Code + ADD
Show the board as a table
measurevanilla+ ADDreading

Which one should your project use?

🏃 Vanilla Claude Code

For throwaway work: a script, a spike, a one-shot. It runs 4–5× faster at about half the price, and on a strong model the code is as good.

🛡️ Claude Code + ADD

For a product you will still be changing next month, with several milestones, teammates, or anything where who may do this? matters. The decisions outlive the chat, the guesses get reviewed, and the checks cannot quietly weaken. Most changes still take the Quick lane: one red→green test and one commit.

Sources: round 6 of the ADD 4.0 pilot, claude-sonnet-5-5 --effort medium, runs vanilla-amb/rep1 and add-4-amb/rep1, condensed from their transcripts (pilot report · results page). "Vanilla" carries the operator's own ~/.claude config, which already asks for red/green TDD. Both workloads are saturated at Sonnet 5.5, and n = 3 shows a direction, not proof.