This is part 2 of a three-part series on harness engineering. Part 1: what it is · Part 2: who checks the work · Part 3: cost and measurement
"Half the time I have to use Codex to verify Claude's work, and bugs still reach PR review anyway."
I typed that into a session in August, mostly venting. Then I went back through about 25 merged pull requests to check.
The bug class nobody catches
I pulled out every logic or contract bug that had got past development, on-device testing and my own AI review, and was caught by a human reviewer or by users. There were thirteen. Some of them:
- A
Booleanin a network model that should have beenBoolean?. It breaks the day the backend leaves the field out. - A periodic background job registered with
KEEPinstead ofUPDATE. Fine on a fresh install, but existing users never get the new schedule. - Logic keyed on
::class.simpleName. Fine in debug, broken in release once R8 renames the class. - A fallback branch deleted in a refactor because it "looked unused." It was load-bearing.
- A precedence chain written backwards.
None of them crash in a demo, and none show up when the agent that wrote the code tests the case it had in mind.
I measured my reviewer
So I built a replay. 450 review threads across 24 PRs. I took 45 bugs that human reviewers had flagged, rolled the code back to the commit before each comment, and ran my AI review setups against it blind. Then I counted.
| Setup | Human-found bugs caught |
|---|---|
| Same review prompt, run one | 21 of 45 |
| Same review prompt, run two | 14 of 45 |
| Prompt plus a list of "things we usually miss" | Lower, every run |
| Strongest model as reviewer, tiered pipeline | 26 of 45 |
| Same pipeline, mid-tier model as reviewer | 19 of 45 |
| Nine independent runs, findings combined | 30 of 45 |
Same prompt, same code, and the two runs were seven bugs apart. That gap was bigger than the effect of almost any prompt change I tried, so a lot of my prompt tuning had been noise, inside the error bars.

Three things did move it:
- More reviewers. Different runs and different tools catch different things. Combining three tools took recall to 51%.
- A stronger model where it matters. A cheap model is fine for sorting files into bundles. It is not fine as the reviewer, and it was worse as the validator: it threw away 54% of real findings, against 14% for the stronger model.
- Smaller pull requests. On the 130 to 177 file PRs, recall collapsed for every setup. Changing the prompt didn't get that back.
Make it a gate
Like the rules in Part 1, "get a second review" only started happening once it was a hook.
Before Claude Code is allowed to run gh pr create, a script checks for a review verdict for this exact diff: branch, merge-base and head commit all baked into the filename. Push one more commit and the old verdict no longer counts. The review itself is done by a different model family with only the diff in front of it. It never sees the plan, the task description or the author's reasoning. So it has to work out from the code alone whether the change is right.

I had Codex review the first version of that gate. It found it trivially bypassable: loose command matching, a skip for detached HEAD, and it quietly allowed everything if jq wasn't installed. All of those paths now block the PR instead of letting it through.
28 tasks done, 3 never built
The one that still bothers me: on a feature built through my spec workflow, the agent marked 28 tasks complete in a single pass using a regex. The build, lint, all 3,460 unit tests and my plan checker passed. Three of those 28 tasks had never been implemented.

The only thing that caught it was a final check that asks, for every task marked done, where is the diff that did it? The other checks only look at code that exists. None of them asked where the code for each ticked task was.
Same pattern at a smaller scale: a checker from the same model family passed a feature, and a checker from a different family blocked it on a real bug. I don't trust a same-family review on its own anymore.
Read the UI tree first
For UI, "verify" used to mean the agent takes a screenshot and says it looks right. Vision models are poor at exactly the things UI bugs are made of: counts, exact colours, whether something is off by a few dp. So the order in my setup is fixed: check the app is running, read the UI tree as text (about 1KB), and only take a screenshot when a human needs to see it. That order is the core of ComposeProof.
What you can do this week
- Replay your reviewer. Take ten bugs your human reviewers caught last month. Run your AI review on the code before each comment. Count. Then run it again and count again. The gap between the two runs tells you how much of any prompt change is noise.
- Never let the authoring session review its own diff. Give it a fresh context and only the diff, ideally on a different model.
- Key review verdicts to the commit, so a review of yesterday's code can't approve today's.
- Add one claim-versus-evidence check: every "done" needs a diff behind it.
- Cap PR size before you tune a single review prompt.
I no longer let the session that wrote the code be the only one that reviews it.
Part 3 is about cost: where my tokens actually went, which model I use for which job, and what I'm starting to measure.