Determinism is a budget, not a property
You cannot make an LLM deterministic. You can decide exactly how much nondeterminism each part of your system is allowed to spend, and enforce it at the boundaries.
"Can we make it deterministic?" is the wrong question, asked in good faith, by someone who has correctly sensed that something is wrong.
The right question: where in this system is nondeterminism acceptable, and what is the blast radius when it shows up?
Draw the line at the boundary
In a well-structured AI feature there are exactly two kinds of component: those allowed to be surprising, and those not. The model is surprising. Everything downstream of it should not be.
That means the model's output crosses a boundary where it stops being prose and becomes data — validated, typed, and rejected if it does not conform.
from pydantic import BaseModel, Field, ValidationError
class Extraction(BaseModel):
invoice_total: float = Field(ge=0)
currency: str = Field(pattern=r"^[A-Z]{3}$")
line_items: list[str] = Field(min_length=1)
def parse(raw: str) -> Extraction | None:
try:
return Extraction.model_validate_json(raw)
except ValidationError:
return None # a caller decides: retry, fall back, or escalate
Past that boundary, everything is ordinary software again, and every ordinary technique applies.
Spend the budget deliberately
Give each surface an explicit allowance:
- Zero. Money movement, permissions, deletions. The model may propose; deterministic code decides.
- Bounded. Classification into a closed set. Surprising output is possible but the range is enumerable, so you can test every branch.
- Open. Summaries, drafts, explanations. Cheap to be wrong, human in the loop.
Most production incidents I have seen in AI features come from a zero-budget operation quietly being handed an open-budget component, usually because a tool was added to an agent without anyone re-asking the question.
The old instinct, again
This is the same discipline as validating input at a trust boundary. We have known it since the first form post. The only thing that changed is that the untrusted input now comes from a component we built ourselves, which makes it feel trustworthy in a way it has not earned.
Subscribe
The weekly essay, free.
One piece a week on what holds up in production. No digest, no roundup, no course funnel.
Related
What a trading engine taught me about agent retries
Exchange infrastructure solved idempotency, retry storms, and partial failure decades ago. Agent frameworks are rediscovering all three the hard way.
The eval harness is the product
If you cannot prove your AI feature got better this week, you are not engineering it — you are decorating it. A working harness, and the failure modes that make most of them lie.