This is part 3 of a three-part series on harness engineering. Part 1: what it is · Part 2: who checks the work · Part 3: cost and measurement
Most arguments I see about AI coding tools are about the model: which one is smarter this month, which one writes better Kotlin. I've stopped measuring the model on its own. What I care about is what a whole run costs and catches, including the context, tools and checks around it.
Where the tokens actually go
In April I wrote about taming my token usage: a session that spawned fourteen identical subagents and burned 20.5M tokens, cut to 2.4M by giving the agent a map instead of letting it grep. That fix held, but it was one session. The totals told me more.
Claude Code keeps a stats file locally. Over the three months it covers, my output was about 59 million tokens. Tokens read back from cache: about 14.2 billion. Roughly 240 tokens re-read for every one written.
Most of what I pay for is the agent re-reading context, not writing.

So output length barely matters for me. What matters is how big the context is every turn, and whether the front of it stays the same so the cache can do its job. Every time a session reorganises its system prompt, swaps tools, or drags a long conversation forward, you pay for it again.
Model choice is per job
The review replay from Part 2 showed where the expensive model is worth it and where it isn't:
| Job | What worked | Why |
|---|---|---|
| Sorting, routing, picking from a list | Smallest model | Cheap, fast, and wrong answers are obvious |
| Writing code | Mid tier, escalate on retry | Most tokens are spent here |
| Judging (review, root cause) | Strongest model | 26 vs 19 caught bugs |
| Validating findings | Strongest model | Mid tier threw away 54% of real findings |
| Checking finished work | A different model family | Same family shares blind spots |

These days I choose a model for each job instead of one for the whole project.
Thirteen rounds of review
More checking costs money, and past a point it stops helping. On one internal tool, my plan checker took thirteen rounds to pass. Rounds three through ten were all the same subsystem: an auto-close feature for duplicate tickets. Every round, the checker found a new edge case. Every round, I hardened it. At round ten I deleted the feature. Duplicates now get flagged and a human closes them.
Now, if the same subsystem fails review three rounds in a row, I change the design.
Plumbing bugs
The failures that cost me the most time this year were plumbing:
claude -pinside awhile readloop swallowed the loop's stdin. Every item after the first was silently skipped.cmd | head -n 5underset -o pipefailexits 141 whenheadcloses early. The step "failed" after succeeding.- A GitHub Actions expression over 21,000 characters makes the whole workflow invalid, silently. My review bot stopped running and nothing told me.
- Two copies of my Slack bot were running at once and splitting events between them. Fixed with a lock file.

I found every one of these the first time the thing ran on real data, not while planning it. Now I run the orchestration under bash -e against real inputs before I trust it.
What I'm starting to track
Claude Code can export its own usage over OpenTelemetry: tokens by type and model, cost, lines changed, commits. On top of that, I'm adding a span for each stage and each gate, the same way I'd trace a network request through an app, so one run reads as a waterfall:
| Area | Tracked |
|---|---|
| Cost | Tokens by type, cache hit rate, cost per stage, subagents spawned |
| Speed | Time per stage, edit-to-green time, time waiting on me |
| Efficiency | Retries per gate, rounds per check, interventions per feature |
| Code | Complexity per function and its change per PR, coverage change |
| Review | PR size, rounds to merge, findings by source, bugs that escaped |
What I want from it: when I change a prompt, a model or a gate, I want to know if things actually improved or if I just had a lucky run.
What you can do this week
- Measure first. Run a token tracker over
~/.claude/projectsand look at cache reads against output. Mine was about 240 to 1. - Keep the front of your context stable. Same system prompt, same tools, short CLAUDE.md. Move reference material to files the agent reads on demand.
- Assign models by job, and write down why, so you can test the choice later.
- Cap verification rounds. If round three finds a new hole in the same place, change the design.
- Test the plumbing with real data before you trust the pipeline.
Seven months in, most of what went wrong for me had little to do with the model itself. It came down to permissions, review, and proof that work was actually done.
I'm putting these hooks, review gates and measurements together into one setup I can reuse. I'll write about it once it works.