Skip to content
Load-Bearing CodeSubscribe
August 13, 2026619 words3 minMembers

The eval harness is the product

If you cannot prove your AI feature got better this week, you are not engineering it — you are decorating it. A working harness, and the failure modes that make most of them lie.


Every team I talk to has the same conversation about six weeks in. Someone changed a prompt, the demo feels better, and nobody can say whether it is better. The honest answer — "we do not know" — is uncomfortable enough that teams stop asking.

The uncomfortable answer is the correct one, and the fix is not a vibe check with more people in the room.

Why the usual approach fails

A twenty-case spreadsheet feels like rigor. It is not. Twenty cases has enough variance that a genuine five-point regression is invisible and a genuine improvement is indistinguishable from a lucky sample.

Worse: the cases are almost always the ones someone remembered, which means they over-represent dramatic failures and under-represent the boring middle where your users actually live.

← All writing