Both runs got a booking-service spec that contradicts itself (a conflict is 202
waitlisted and 409 rejected) and says nothing on six other decisions. Watch where each
one puts its decisions, and what that costs.
Both runs built working code, and both caught the contradiction. The difference is where each decision lives once the chat is gone, and the guess nobody asked about.
"waitlist": false opt-out, in its chat reply.R:OWNER, pinned by sealed check C11.freeze, build, verify.One pair of runs can mislead, so the board shows means over n = 3 runs per arm, on both workloads: wm1, a booking API built from scratch, and amb1, the contradictory spec above.
| measure | vanilla | + ADD | reading |
|---|
For throwaway work: a script, a spike, a one-shot. It runs 4–5× faster at about half the price, and on a strong model the code is as good.
For a product you will still be changing next month, with several milestones, teammates, or anything where who may do this? matters. The decisions outlive the chat, the guesses get reviewed, and the checks cannot quietly weaken. Most changes still take the Quick lane: one red→green test and one commit.
Sources: round 6 of the ADD 4.0 pilot, claude-sonnet-5-5 --effort medium,
runs vanilla-amb/rep1 and add-4-amb/rep1, condensed from their transcripts
(pilot report ·
results page).
"Vanilla" carries the operator's own ~/.claude config, which already asks for red/green TDD. Both
workloads are saturated at Sonnet 5.5, and n = 3 shows a direction, not proof.