Evaluation and testing / ENGINEERING NOTE
How to test an LLM workflow when the answer changes every time
An LLM generates this test description:
A €600 expense with a receipt requires neither manager approval nor a receipt request.
The next run produces:
For this receipted €600 claim, no manager sign-off or receipt follow-up is needed.
A string comparison rejects the second answer. A reviewer can accept both.
Now consider a third answer:
A €600 expense must receive manager approval.
It sounds reasonable. It uses the same vocabulary. It is wrong for our fictional policy, which requires manager approval only above €750.
I want tests that tolerate changes in expression while exposing changes in behavior. That starts with defining what an output must mean, what the system can check, and what still needs a person’s judgement.
In the previous article, I made an LLM workflow survive a crash without discarding saved progress. Recovery preserves outputs; it does not establish whether they are worth preserving. This article addresses that next problem.
Start with the contract, not the reference paragraph
I use the same fictional expense-policy application, with two requirements:
- Expenses strictly above €750 require manager approval.
- Every expense requires a receipt. If it is absent, request it.
The generator produces a suite of structured test cases, each with a description for reviewers. For this small example, the suite must exercise an amount below the threshold, exactly at it, above it, and an expense without a receipt.
Those are evaluation requirements I specify before looking at a candidate answer. Four fluent descriptions do not necessarily satisfy them.
A generated case looks like this:
{
"id": "T-600",
"amountCents": 60000,
"hasReceipt": true,
"expected": {
"approvalRequired": false,
"receiptRequired": false
},
"sourceIds": ["approval", "receipt"],
"description": "A €600 expense with a receipt requires neither manager approval nor a receipt request."
}
Here, receiptRequired means that this test expects the application to request a missing receipt. It does not mean that the underlying policy stops requiring receipts when one is already attached.
The containing suite also records sourceVersion: "expense-v2". A correct answer to yesterday’s policy is not a correct answer to today’s request.
The IDs, booleans, amounts, and source references make several properties testable without interpreting prose. The description remains a separate part of the product that someone must assess.
Use several checks with different responsibilities
I split evaluation into layers so that a failure tells me what to investigate:
| Layer | Question | Example failure |
|---|---|---|
| Parsing | Can I read the response? | Truncated JSON |
| Structure | Does it satisfy the output contract? | A string where a boolean is required |
| Provenance and identity | Is it tied to the intended source and unambiguous records? | Old source version or duplicate test ID |
| Correctness | Do structured expectations follow the policy? | Approval required for €600 |
| Coverage | Are the required scenarios actually exercised? | No missing-receipt case |
| Review | Does the explanation agree with the assertions and avoid invented requirements? | Prose adds director approval |
A parsing failure stops subsequent checks. There is no meaningful coverage score for a payload I cannot inspect. A structurally valid payload proceeds to the other deterministic checks, which can report several failures together.
JSON Schema can express types, required fields, and whether additional properties are allowed. These are useful contract checks; they do not by themselves establish that an approval decision follows the source policy. JSON Schema object reference
The reference implementation uses a small explicit validator rather than a schema library. It checks exact field names, safe non-negative integer cents, booleans, non-empty descriptions, and a maximum of 100 cases. It does not coerce "false" into false.
For a production interface, I would normally keep the schema versioned and validate response size before parsing. The example’s array limit is not a general protection against oversized untrusted input.
Write the business oracle independently
The generated answer must not supply its own definition of correctness.
For this fixture, I encode the trusted policy directly:
const approval = test.amountCents > 75000;
const receipt = !test.hasReceipt;
if (test.expected.approvalRequired !== approval) {
fail('correctness', 'wrong_approval');
}
if (test.expected.receiptRequired !== receipt) {
fail('correctness', 'wrong_receipt');
}
Integer cents avoid introducing floating-point ambiguity into this boundary. The operator matters: > and >= disagree at exactly €750.
This is a domain oracle for one controlled policy. It does not extract arbitrary requirements from documents. I can trust it only after checking its interpretation against the requirement, including that exact boundary.
The fixtures contain independently written expected outcomes, including a deliberate boundary error. I do not create all expected answers by calling the same helper that I am testing. If both generator and grader inherit the same mistaken interpretation, a green report becomes misleading agreement.
In a real project, I would have a domain owner confirm the policy interpretation. Disagreement about what the requirement means needs resolution before it can become a useful regression test.
Measure coverage using inputs and assertions
Suppose a generated suite contains four versions of the €600 test, all with correct expectations. It contains four cases, but it still misses the threshold and the missing receipt.
I derive coverage from the actual input values and their validated expectations. I do not trust a model-supplied label such as "boundary test".
In this example:
amountCents < 75000can cover the below-threshold category.amountCents === 75000can cover the boundary.amountCents > 75000can cover the above-threshold category.hasReceipt === falsecan cover the missing-receipt category.
A case contributes only if both structured expectations are correct. A €750 case with the wrong approval result cannot make the boundary category green.
Other defects still reject the suite. A case might exercise the numerical boundary while citing an unknown source; coverage does not cancel the provenance error.
This is intentionally a minimum coverage contract. It does not establish coverage of every combination. If the product needs above-threshold expenses without receipts, multiple currencies, or role-dependent exceptions, those deserve explicit requirements and cases. I would expand the contract rather than infer completeness from a percentage.
Keep prose quality separate from machine-checkable correctness
The most instructive fixture has correct booleans and an incorrect description:
amountCents: 60000
approvalRequired: false
description: A €600 expense must receive manager approval.
The numerical oracle passes. The description contradicts it.
The evaluator deliberately does not pretend to understand that sentence. It returns:
{
"deterministic": "pass",
"decision": "review",
"reviewRequired": [
"description_assertion_consistency",
"unsupported_narrative_claims",
"reviewer_clarity"
]
}
That is an abbreviated report; the full output also includes failures and coverage. A deterministic pass qualifies the suite for review. It is not an approval decision.
The same boundary applies to hallucinated requirements. An unknown source ID is easy to reject against an allowlist. A sentence that invents director approval while citing two valid source IDs requires a content review. Valid citations do not prove that the cited material supports every claim.
If the product needs only structured fields, I would consider rendering descriptions from those fields with a template. That removes one source of disagreement. If generated explanations are useful, I need to evaluate them as outputs in their own right.
Give a model judge a narrow job
A model judge could help review the descriptions, but I would not replace all these checks with “rate this answer from one to ten.”
My proposed rubric has three questions:
| Dimension | Pass condition | Failure example |
|---|---|---|
| Consistency | The description agrees with the inputs and expected results | Prose requires approval; boolean says it is unnecessary |
| Support | Every claimed requirement is supported by the supplied policy | Director approval appears without a source requirement |
| Clarity | A reviewer can understand what is being exercised | “Check that it works correctly” |
I would supply the policy, the candidate case, and those criteria. The judge should return one assessment per dimension, identify the relevant words, and allow an uncertain result. Candidate text is material to inspect, never authority to change the rubric.
Before trusting that judge, I would compare its decisions with independent human reviews, including plausible but incorrect answers. I would look separately at wrong answers accepted and good answers rejected. A high overall agreement figure can hide a judge that accepts the particular defect I most need it to catch.
Anthropic’s evaluation guidance distinguishes code-based, model-based, and human graders and recommends calibrating model graders against human judgement. I use that distinction here to assign each checker a bounded responsibility. Demystifying evals for AI agents
The runnable example contains no model judge and makes no claim about judge accuracy. Its review notes are authored fixture labels, not the result of an independent human review study. The contradictory-description fixture shows a deterministic checker’s limit; it is not a measured judge-versus-human disagreement.
Test your evaluator before comparing models
The companion example contains 13 hand-authored candidate suites: seven development fixtures and six reserved regression fixtures. Some are deliberately wrong. They are test inputs for the grader, not outputs sampled from an LLM.
| Candidate | Deterministic result | Next action |
|---|---|---|
| Valid original description | Pass | Review prose |
| Valid paraphrase | Pass | Review prose |
| €600 expectation using the old threshold | Fail | Reject |
| Missing receipt scenario omitted | Fail | Reject |
| Invented source ID | Fail | Reject |
| Truncated JSON | Fail | Reject |
| Correct fields, contradictory description | Pass | Review; authored label identifies the contradiction |
The reserved fixtures include an exact-boundary error, wrong receipt outcome, duplicate ID, stale version, invented narrative requirement, and reordered valid suite. Separating them illustrates how to keep development examples apart from later checks. Since they are small fixtures authored for this demonstration, this is not an independently collected held-out benchmark.
The test suite also checks that malformed inputs stop safely, values are not coerced, repeated correct cases cannot hide missing coverage, and deterministic passes never silently approve prose.
Download the reference example ZIP, extract it, and run these commands in its folder:
node --test evaluate.test.mjs
node run.mjs
node run.mjs --reserved
The reference example runs with Node.js and needs no additional dependencies or API credentials. It reports individual failures rather than collapsing everything into a quality score.
Its 18 passing tests establish the declared behavior of this evaluator. They do not mean an LLM passed 18 tasks, and rerunning deterministic fixtures does not measure model variability.
Compare changes over repeated trials
Once the grader is useful, I can use it to compare actual generation runs.
I would preserve the input fixture version, source snapshot, prompt revision, model identifier, generation settings, grader revision, and raw output for every trial. If the input, prompt, and grader all change together, I cannot confidently attribute the difference to one cause.
For a prompt change, I would run the baseline and candidate against the same fixture set, multiple times, with isolated workflow state. I would keep development examples separate from the evaluation set and avoid repeatedly tuning against supposed held-out cases. Once I use a case to revise the prompt, it has become development feedback.
My comparison would show:
- Valid responses out of attempted trials, with infrastructure failures identified separately.
- Correctness failures by rule and scenario, especially boundaries and exceptions.
- Missing required coverage by suite.
- Human or calibrated-judge decisions, including unresolved disagreements.
- Latency, token use, and retry counts alongside quality.
I would inspect regressions case by case before accepting an improved average. Fewer malformed responses would not compensate for a newly introduced approval error in this application.
The number of trials should reflect the failure frequency and consequence I need to detect. A handful of successful runs cannot establish that a rare failure is absent. I would report sample counts and the limits of the evidence rather than label a small experiment “production accuracy.”
No repeated model experiment has been run for this article. This section describes how I would extend the fixture runner into one.
Make the release decision explicit
For this example, the decision rule is simple:
Parsing or contract failure → reject.
Wrong policy outcome, invalid provenance, or missing required coverage → reject.
All deterministic checks pass → review the remaining content.
That is a policy for this demonstration, not a universal threshold for LLM applications. A system drafting suggestions for a reviewer can accept a different level of uncertainty from one taking actions automatically.
I would also test the surrounding workflow: which source reached the model, whether a failure triggers an allowed retry, whether stale results are rejected, and whether a review decision attaches to the output actually reviewed. A correct grader cannot compensate for evaluating one output and publishing another.
The useful shift is to stop asking whether the answer matches one paragraph and ask whether it satisfies an explicit contract. Stable wording is optional. Correct behavior, relevant coverage, and an honest account of what remains unchecked are not.
For the broader project context, see SESCA. For how source changes determine which saved outputs need another look, start with selective regeneration in LLM workflows.