The eval harness is the product
If you cannot prove your AI feature got better this week, you are not engineering it — you are decorating it. A working harness, and the failure modes that make most of them lie.
Every team I talk to has the same conversation about six weeks in. Someone changed a prompt, the demo feels better, and nobody can say whether it is better. The honest answer — "we do not know" — is uncomfortable enough that teams stop asking.
The uncomfortable answer is the correct one, and the fix is not a vibe check with more people in the room.
Why the usual approach fails
A twenty-case spreadsheet feels like rigor. It is not. Twenty cases has enough variance that a genuine five-point regression is invisible and a genuine improvement is indistinguishable from a lucky sample.
Worse: the cases are almost always the ones someone remembered, which means they over-represent dramatic failures and under-represent the boring middle where your users actually live.
Related
Determinism is a budget, not a property
You cannot make an LLM deterministic. You can decide exactly how much nondeterminism each part of your system is allowed to spend, and enforce it at the boundaries.