Since early 2026 I've been building my own harness around Claude Code for my day-to-day work on an Android app. Mostly hooks that block things, a second model reviewing the first one, and keeping count of what it costs me. This is what I've learned so far, a little over 24,000 prompts in.
By harness I mean everything around the model that you own: CLAUDE.md, skills, hooks, MCP servers, CI, the review process, and the scripts that hold it together. People started calling this harness engineering earlier this year (Mitchell Hashimoto, OpenAI, and Birgitta Böckeler on Martin Fowler's site all wrote about it). Those posts are mostly by people building agents. I use one every day on an Android app, so that's the part I can write about.
Rules in CLAUDE.md don't hold
In August I asked my agent to run the instrumented tests on my test phone. The Gradle task uninstalled the app first, which wiped its data and logged me out of my test account.
I already had a rule against exactly this. My CLAUDE.md said, in capitals: never clear app data without asking. The agent didn't clear the data itself. It ran a task that did. The rule listed pm clear and uninstall. It didn't list a Gradle test task, because I didn't know that task uninstalls first.
Every time I stop the agent and say "no, not like that", I write it down: what happened, the rule, why. There are 24 of those now. In almost every case the model could do the work. It just did more than I asked. It edited code when I'd asked for analysis, applied a design I'd only pasted as a reference, and installed a build over a call I was still testing.
Every one of those rules started as a sentence in CLAUDE.md, and several came back anyway. The fully qualified names rule took two nudges on consecutive days. The rules that stopped coming back were the ones I moved into code:
- A hook that runs before every tool call and checks each edit against my folder of past mistakes, each one a file glob plus a regex. Claude Code blocks the call when the hook exits with code 2.
- A read-only mode for questions I ask from Slack, where the agent simply doesn't get write tools.
- A guard that knows when another session is using my test phone and refuses an install on it.
- An allowlist instead of a denylist for my Slack bot's commands. The denylist leaked:
phpcould run arbitrary code andgit update-refcould move branches.

Name the effect, not the command, and enforce it in code.
Who checks the agent's work
"Half the time I have to use Codex to verify Claude's work, and bugs still reach PR review anyway." I typed that into a session in August, then went back through about 25 merged pull requests to check. I pulled out every logic or contract bug that had got past development, on-device testing and my own AI review, and was caught by a human reviewer or by users. There were thirteen. Some of them:
- A
Booleanin a network model that should have beenBoolean?. It breaks the day the backend leaves the field out. - A periodic background job registered with
KEEPinstead ofUPDATE. Fine on a fresh install, but existing users never get the new schedule. - Logic keyed on
::class.simpleName, which R8 renames in release. - A fallback branch deleted in a refactor because it looked unused, though the code still depended on it.
None of these crash in a demo, and the agent that wrote the code misses them when it only tests the case it had in mind.
So I replayed my AI review against 45 bugs that human reviewers had flagged on 4 PRs, rolling the code back to the commit before each comment.
| Setup | Human-found bugs caught |
|---|---|
| Same review prompt, run one | 21 of 45 |
| Same review prompt, run two | 14 of 45 |
| Strongest model as reviewer | 26 of 45 |
| Mid-tier model as reviewer | 19 of 45 |
| Nine independent runs, findings combined | 30 of 45 |

Two runs of the same prompt on the same code were seven bugs apart, and most of my prompt changes moved the result by less than that. Combining independent reviewers helped. So did using the stronger model to review and to validate findings: as validator, the mid tier threw away 54% of real findings and the stronger model 14%. Smaller pull requests helped too. On the 130 to 177 file PRs every setup did badly.
Like the rules above, "get a second review" only started happening once it was a hook. Before Claude Code can run gh pr create, a script checks for a review verdict tied to the exact branch, merge-base and head commit, so a new commit needs a new review. The reviewer is a different model family and sees only the diff. Codex reviewed the first version of that gate and found it trivially bypassable, so now it fails closed.
The one that still bothers me: on one feature the agent ticked 28 tasks done in a single pass. The build, lint, all 3,460 unit tests and my plan checker passed. Three of those 28 tasks had never been implemented. The only check that caught it asked, for each ticked task, where the diff was.

What it costs
In April I wrote about taming my token usage: one session spawned fourteen identical subagents and burned 20.5M tokens, cut to 2.4M by giving the agent a code index instead of letting it grep. Then I looked at the totals. Over the three months Claude Code's local stats cover, it wrote about 59 million tokens for me and re-read about 14.2 billion from cache. Roughly 240 re-read for every one written.

For my usage, output is a small part of the bill. Most of it is context re-read on every turn, which is also why I try to keep the start of that context unchanged so the cache keeps working. From the review tests, this is how I split the jobs:
| Job | What worked |
|---|---|
| Sorting, routing, picking from a list | Smallest model |
| Writing code | Mid tier, escalate on retry |
| Review and validation | Strongest model |
| Checking finished work | A different model family |

More checking has a cost too. On one internal tool my plan checker took thirteen rounds to pass, and rounds three through ten all found new edge cases in the same auto-close feature. At round ten I deleted the feature. Now, if the same subsystem fails review three rounds in a row, I change the design.
The failures that cost me the most time this year came from my own scripts and workflows. claude -p inside a while read loop swallowed the loop's stdin, so every item after the first was silently skipped. A GitHub Actions expression over 21,000 characters made a whole workflow invalid without telling anyone, and my review bot quietly stopped running.
What I'd do this week
- Start a corrections file. Every "no, not like that" gets a line: what happened, the rule, why.
- Move your worst CLAUDE.md rule into a hook. Have it check the effect, for example anything that clears app data.
- Replay your reviewer. Take ten bugs your human reviewers caught last month and run your AI review on the code before each comment, twice. Running it twice shows how much the result moves without touching the prompt.
- Never let the session that wrote the code be the only one reviewing it.
- Look at your cache reads against output in
~/.claude/projects. Mine was about 240 to 1.
Next, I'm moving all of this onto my own setup built on OpenCode, with OpenRouter for models. I'll write that up once it's running.