WORK WITH MEOpen for new builds · 2026
← BLOG

The real AI bill is cost per successful workflow

Token prices matter. Retries, review effort and failed work determine whether the system is economical.

A branded cost-per-success illustration on a black background, titled Measure cost per success.

When an AI pipeline looks expensive, the first reaction is often to choose a cheaper model. Sometimes that helps. Sometimes it shifts the bill into retries, manual correction and support. I want the business to see the cost of completing useful work, including the cases that go wrong.

Choose a useful unit of cost

My preferred unit is cost per accepted outcome. For a defined workload, add model usage, tools, infrastructure, review effort and operating support, then divide by outcomes that meet the acceptance standard. Keep one-off development investment visible separately or amortise it explicitly. Do not hide it in a mysteriously low operating number.

Consider a deliberately simplified illustration. A batch costs €12 in model and tool usage and €48 in review time. If 80 outcomes are accepted, that is €0.75 per accepted outcome. A second design costs €20 in usage and €20 in review, with the same accepted volume: €0.50. These are invented teaching numbers, not provider prices or client results. They show why the cheapest inference can produce the more expensive service.

Reduce waste in the workflow

The harness gives us places to improve. Repeated full-document reads, oversized tool results, unnecessary model turns and loops without a useful stopping condition are candidates for investigation. A trace should reveal which step consumed the budget and whether that step improved the result.

Context engineering means choosing relevant information rather than accumulating everything available. Anthropic’s guidance describes context as a resource to curate throughout an agent’s run. Anthropic — Effective context engineering for AI agents Its tool-use work also shows why discovering tools when needed and processing intermediate results outside the model can reduce avoidable context load. Anthropic — Introducing advanced tool use Neither technique should be adopted without checking its effect on the actual workload.

For a proposed support workflow, I would retrieve the relevant policy section, keep a reference to the source, and perform straightforward lookups in code. I would test whether a smaller model handles classification, while reserving more capable inference for cases that need it. The escalation rule would come from observed performance, not the model’s self-reported confidence alone.

Put limits around each run

Hard budgets matter. Limit steps, elapsed time, tool calls and spend per run. When a case reaches its limit, stop with a useful handover containing what was tried and what remains unresolved. An agent repeatedly attempting the same action is generating activity, not necessarily value.

Finally, look at the distribution. An attractive average can conceal a small group of extremely expensive cases. Track the tail, review workload, unsuccessful runs and latency alongside quality. The business decision is whether this complete service improves the work at an acceptable cost. Token optimisation earns its place when that answer becomes stronger.


For a practical review of workflow quality, cost and bottlenecks, contact hi@fdo.codes.

(Contact)

LET'S
BUILD.