Reliability means passing the same standard, not repeating the same sentence
Evaluate outcomes, exceptions and actions across repeated runs before expanding autonomy.

Two AI answers can be worded differently and both be correct. Two identical answers can both be wrong. That is why I would not sell “uniform results every time” as the promise of a model harness. I would define a standard the system must meet and measure how consistently it meets it.
The standard depends on the job. A meeting-summary tool must represent decisions accurately and distinguish agreed actions from suggestions. A support assistant must use the applicable policy and avoid unsupported commitments. A coding agent must produce a change that works in the target environment without breaking required behaviour.
Test the difficult cases
An evaluation set should contain ordinary work and the cases people find difficult. For an illustrative meeting-to-action pipeline, I would include contradictory statements, a changed deadline, an absent owner, a cancelled task and a transcript containing instructions aimed at the agent. The expected behaviour may be a correct result, an explicit uncertainty or a decision to ask a person.
I would test representative cases multiple times. One successful demonstration is weak evidence about a probabilistic system. Anthropic’s evaluation guidance separates a task from an individual trial and distinguishes the agent’s transcript from the actual outcome in its environment. Anthropic — Demystifying evals for AI agents That matters whenever the system can act: claiming a task was created is different from a verified task in the right project.
Combine checks with clear release criteria
My evaluation plan would combine several checks. Code can verify a required schema, permitted identifiers and observable system state. Human reviewers can judge usefulness against a written rubric. Model-based grading can help at scale, but needs calibration against human judgement; another model is not an infallible judge.
Agree the release threshold before examining the latest results. Which errors are tolerable? Which actions are prohibited? How often may the system escalate? What latency and review burden can the team accept? There is no universal percentage that answers these questions for every workflow.
Test the harness as well as the answer. Did the run respect tool permissions? Did it stay within its budget? Did it preserve evidence for review? Did a failed dependency trigger a useful fallback? A fluent final paragraph can conceal a poor execution path.
Keep learning after release
After release, sample real outcomes and maintain a regression set of meaningful failures. Refresh cases as the work changes. Preserve a held-out set where possible so the team does not merely optimise for familiar examples.
The client should see the system’s strengths and limits in terms of their work. That is a useful form of confidence: success is testable, failure is visible, and the team knows what to do when a case falls outside the expected boundary.
I help teams define acceptance criteria and build evaluations into the delivery workflow. hi@fdo.codes







