This is part 1 of a three-part series on harness engineering. Part 1: what it is · Part 2: who checks the work · Part 3: cost and measurement
In August I asked my agent to run the instrumented tests on my test phone. It did. The Gradle task uninstalled the app as part of the run, which wiped the app's data, which logged me out of the test account I was using.
I already had a rule for this. It was in my CLAUDE.md, written in capitals after the last time: never clear app data without asking. The agent had read it. It didn't clear the data. It ran a task that cleared the data. It followed the rule, and I still lost the data.
24,000 prompts later
Since April I've sent Claude Code a little over 24,000 prompts. Most at my day job on a large production Android app, the rest at night on side projects. Back in March I wrote about the correction tax and said the skill that matters now is building the harness around the AI. Seven months on, I can be more specific about what I meant.
The harness is everything around the model that you own: CLAUDE.md, skills, hooks, MCP servers, CI, the review process, and the shell scripts holding it together. The industry started calling this harness engineering earlier this year (Mitchell Hashimoto, OpenAI, and Birgitta Böckeler on Martin Fowler's site all wrote it up). Most of that writing comes from people who build agents. I use one every day to ship a mobile app, and I've kept the numbers.
The corrections file
Every time I stop the agent and say "no, not like that," I write it down: what happened, the rule, why. There are 24 of those now.
| What the agent did | Category |
|---|---|
| Edited code while I'd asked for analysis only, on the wrong branch | Did too much |
| Applied a design I'd pasted for reference, not for implementation | Did too much |
| Installed a build over a live test call I was in the middle of | Did too much |
| Ran a test task that uninstalled the app and wiped my login | Did too much |
| Wrote a confident root cause off a too-narrow crash window | Claimed too much |
| Inlined fully qualified names, again, after being told | Ignored a rule |
Almost none of the 24 are about the model not being able to do something. Most are about it doing more than I asked for.
Prose rules come back
Every one of those rules started as a sentence in a markdown file the agent loads at the start of every session. Several came back anyway. The fully qualified names one took two nudges on consecutive days. The no-blockquotes one has a "repeat violation" note in its own file.

And then there's the uninstall. The rule said "never clear data." It listed the commands: pm clear, uninstall. It didn't list a Gradle test task, because I didn't know that task uninstalls first. So the rule covered the commands I knew about and missed the one I didn't.
Name the effect, not the command, and enforce it in code.

What held
The rules that stopped coming back were the ones I moved out of CLAUDE.md and into code.
- Claude Code runs a script before every tool call. Exit with code 2 and the call doesn't happen; the agent reads your message instead. My write-time hook checks every edit against a folder of past mistakes, each one a file glob plus a regex. If an edit matches a mistake I've already made once, it's blocked before it lands.
- When I run Claude Code from Slack in read-only mode, I don't tell it to be careful. I don't give it write tools. The setup refuses to finish if a test write succeeds.
- I run several sessions at once against one phone. One session has no way to know another is mid-test. A small extension watches every device command across sessions and blocks an install on a phone someone else used in the last fifteen minutes. The model doesn't need to know the other session exists.
- My first permission list for the Slack bot blocked dangerous commands by name. It leaked:
phpcould run arbitrary code,git update-refcould move branches. I flipped it to allow only what's needed and re-check it on every run.
| Rule in CLAUDE.md | Rule as a hook or mode | |
|---|---|---|
| Agent can ignore it | Yes | No |
| Agent can find a way around it | Yes, by doing the same thing another way | Only if you named the wrong thing |
| Costs context every session | Yes | No |
| Tells you when it fires | No | Yes |
The two-minute test
Open your CLAUDE.md. Find every rule you've written in capitals, or written twice, or written after an incident. For each one, ask: if the agent ignored this sentence tomorrow, what would stop it?
If nothing would, that rule isn't doing much. Pick the one that would hurt most and turn it into a hook this week. In my setup that's usually about twenty lines of shell.
What you can do this week
- Start a corrections file. Every "no, not like that" gets a line: what happened, the rule, why.
- Move your worst rule into a PreToolUse hook. Have it check the effect, for example anything that clears app data.
- Split read from write. If a task is investigation, run it without write tools at all.
- Flip any denylist to an allowlist. It's a much shorter list to get right.
Most of my 24 corrections were about limits, and the limits only stuck once I put them in code.
In Part 2 I replay my AI code review against 45 bugs that human reviewers caught, and look at why the session that wrote the code shouldn't be the one reviewing it.