BLOG
(Notes on building with AI)Agentic AI without the hype: where agents earn their place
An agent is a language model, a set of tools and a loop. Where that combination pays off in real organisations, and where a plain workflow is still the better choice.
“Agent” has become one of the most stretched words in technology. It gets attached to chatbots, scripts with a model call inside, and genuinely autonomous systems alike. That makes it hard to decide where agents belong in an organisation, and easy to build one where something simpler would do better.
Here is the definition I work with: an agent is a language model, a set of tools and a loop. The model looks at the goal and what it knows so far, decides on the next step, calls a tool, reads the result and repeats until the job is done or it hands over to a person. Everything else (memory, planning, multiple agents talking to each other) builds on that loop.
Workflow or agent?
The first question is not which agent to build, but whether you need one at all.
If you can draw the steps in advance, build a workflow: a fixed sequence where a model handles specific steps, such as reading a document or classifying a request, and ordinary code handles the rest. Workflows are predictable, cheap to run and easy to test.
Reach for an agent when the path depends on what you find along the way. Investigating an incident, chasing down why two records don’t match or assembling context from several systems are hard to script, because the next step depends on the last result.
In practice, most production systems I build are workflows with agentic pockets: a deterministic backbone, with an agent handling the steps that genuinely need judgement.
Where agents earn their place
The strongest use cases I see share three traits: high volume, messy inputs and a clear definition of a good outcome.
- Triage and classification. Reports arrive in different formats and languages, and each one needs a category, a priority and an owner. An agent can read, enrich and route them in seconds, so people spend their time on the cases that matter.
- Investigation. Pulling context from case histories, documents and databases to prepare a decision is exactly the kind of multi-step, path-dependent work agents handle well.
- Reconciliation. Matching records across systems and explaining the exceptions turns a periodic manual exercise into something that can run continuously.
- Document review. Flagging documents that look manipulated, incomplete or inconsistent, for a person to check, rather than making the final call.
Notice that none of these remove the human. They change where the human spends their attention.
The parts that matter more than the model
When an agent misbehaves, the model is rarely the root cause. These usually are:
- Tools. Each tool should do one thing, be clearly described and only have the permissions it needs. Standards such as the Model Context Protocol (MCP) make it much easier to expose systems to agents consistently, but the design discipline is still yours.
- Context. An agent can only reason about what it can see. Decide what it retrieves, what it remembers between runs and what it should never see.
- Guardrails. Spell out what the agent may never do, cap how many steps and how much budget it can use, and require approval before anything irreversible.
- Observability. Log every step and tool call. When something goes wrong, you need to replay the run, not guess.
- Evaluation. A set of real cases with known good outcomes, run before every change. Without it, you are tuning by anecdote.
Humans in the loop, on purpose
“Human in the loop” is often added at the end as a safety blanket. It works better when it is designed in from the start, at the points where judgement or accountability actually sits:
- approval before actions that are hard to undo, such as payments, external messages or record changes
- a review queue for low-confidence results, with the agent’s reasoning attached
- a clear escalation path when the agent is stuck, instead of letting it try forever
The goal is not to remove people. It is to move them to the decisions that need them.
The same holds when the agent writes the code
AI coding agents now write much of the code on my client builds. The pattern is the same as in any agentic system: the agent handles volume, and I own the framing, the architecture, the review and the final engineering call. The agent is fast; the direction still has to come from someone who understands the problem and is accountable for the result.
A short checklist before you build one
- Can you write down what a good outcome looks like, and check it on real examples?
- Could a workflow do the job? If so, start there.
- Does each tool have the narrowest permissions that still work?
- Where are the human checkpoints, and who staffs them?
- Can you replay any run from its logs?
- Who owns the system after launch?
If you can answer all six, an agent is very likely worth building. If you can’t, the gaps are where the project would have stalled anyway.
Choosing a language model in 2026: what actually matters
New models arrive every few weeks. The leaderboard rarely decides what works in production; your tasks, your data and your constraints do.
A new language model is released every few weeks, each one topping some leaderboard. It is tempting to treat model choice as the main decision in an AI project and to switch every time the rankings move.
In my experience, the leaderboard rarely decides what works in production. Your tasks, your data and your constraints do. The model is one component of the system, and the system is what has to deliver.
Think in tiers, not names
The major providers offer families of models in roughly three tiers, and they are worth thinking about separately:
- Frontier models are the most capable, and the slowest and most expensive. Use them for complex reasoning, long multi-step agentic tasks and difficult code.
- Mid-tier models handle the bulk of production work well: drafting, summarising, extraction with judgement, most tool use.
- Small, fast models are ideal for classification, routing, simple extraction and anything you run at very high volume.
Anthropic’s Claude models, for example, come as Opus, Sonnet and Haiku; OpenAI and Google offer similar ranges. Alongside them sit open-weight families such as Llama, Mistral and Qwen, which you can host yourself.
Good systems often use several tiers at once: a small model triages and routes, and a larger one handles the cases that need real reasoning. Matching the model to the step usually saves more money than any discount negotiation.
Evaluate on your own work
Public benchmarks measure general ability. They tell you little about how a model handles your documents, your terminology and your edge cases.
What works is a small evaluation set built from real examples:
- Collect 50 to 200 real cases, including the awkward ones.
- Write down what a correct result looks like for each.
- Score each model on accuracy, the kinds of mistakes it makes, whether it follows the required format and how reliably it uses tools.
- Run the same set every time a new model comes out, or whenever you change a prompt.
An afternoon of evaluation beats a week of reading release notes.
Once that set exists, deciding whether to adopt a new model takes hours instead of weeks of debate.
The criteria that actually decide
When I compare models for a production system, these are the questions that matter:
- Quality on your evaluation set. Not in general; on your tasks.
- Tool use and instruction following. For agentic work, the model must call tools correctly, stay within its scope and stop when it should. This varies more between models than raw intelligence does.
- Latency. A model that is a little better but twice as slow can make an interactive product feel broken.
- Cost per completed task. Count the whole task, not the price per token: long prompts, retries and multiple steps add up. Prompt caching and batch processing can change the picture significantly.
- Context handling. A large context window is not the same as using that context well. Test retrieval from long inputs explicitly.
- Data protection. Where is data processed, how long is it retained, is it used for training, and is there a proper data processing agreement? For European organisations, this often narrows the choice before quality is even discussed.
- Reliability. Rate limits, regional availability and how the provider handles outages.
Hosted or open-weight?
Hosted frontier models give you the most capability with the least operational effort. Open-weight models give you control: data never leaves your infrastructure, you can fine-tune, and costs are predictable at scale.
The trade-off is operational effort and, for the hardest tasks, usually some capability. A common and sensible split is hosted models for complex reasoning and open-weight models for high-volume or particularly sensitive steps.
Design for change
The one thing I can predict is that today’s best choice will not be next year’s. So I build systems where switching models is a configuration change followed by an evaluation run, not a rewrite:
- a thin layer between the application and the model provider
- prompts and evaluation sets kept under version control, next to the code
- the model for each step set in configuration, not hard-coded
With that in place, a new release becomes an opportunity rather than a migration project.
Where I start
My default is to start each step with a strong mid-tier model, measure it on real cases, then move the hardest steps up a tier and the high-volume, simple steps down. I re-run the evaluations whenever a significant new model arrives, and only switch when the numbers say so.
The model matters. But the evaluation set, the tool design and the guardrails around it are what turn a good model into a system people can rely on.
Frame, orchestrate, evaluate, ship: how I take AI from idea to production
The four stages I use to turn an AI idea into a system people rely on, and the question each stage has to answer before the next one starts.
Most AI projects that stall don’t fail because of the model. They fail at the edges: a problem that was never sharply defined, no way to tell whether the system is better than what it replaces, or nobody owning it after launch.
To avoid those failures, I take every AI initiative through four stages: frame, orchestrate, evaluate and ship. Each stage ends with a question that has to be answered before the next one starts. Skipping a stage doesn’t save time; it just moves the cost to later, when it is more expensive.
1. Frame
Start with the business problem, not the technology. The questions at this stage:
- What should change, and how will we measure it? Fewer hours spent, faster turnaround, fewer errors, earlier detection. Pick one primary number and record today’s baseline.
- Where is human judgement required? Some decisions should stay with people for reasons of accountability, regulation or trust. Name them now, not after the design is done.
- What data exists, and can we use it? Availability, quality, ownership and data protection rules all shape what is possible.
- What are the risks? Wrong answers, biased outcomes, security, and obligations under rules such as the GDPR and the EU AI Act.
The output is a short brief: the problem, the metric, the baseline, the boundaries and the owner. If that brief can’t be written, the project isn’t ready for a design.
Exit question: do we agree on what success looks like, and how we’ll know?
2. Orchestrate
Now design how it works. This is where the architecture decisions are made:
- Workflow or agent? If the steps are known in advance, a workflow is simpler and more reliable. Agents are for the parts where the path depends on what is found along the way.
- Tools and data access. Which systems does it read from and act on, with which permissions? Expose them as narrow, well-described tools.
- Guardrails. What must it never do, which actions need approval, and what happens when it is unsure?
- Human checkpoints. Where people review, approve or take over, and who those people are.
AI agents are part of how I build these systems too. Coding agents turn the design into working software quickly, while I direct the process, set the architecture and review what they produce.
The output is a working prototype running on real data, not a slide deck.
Exit question: does it work on real cases, end to end?
3. Evaluate
This is the stage most often skipped, and the one that decides whether a system can be trusted.
- Build an evaluation set from real cases with known correct outcomes, including the difficult ones.
- Measure against the baseline from the framing stage, not against an ideal.
- Study the failures, not just the score. Which kinds of cases go wrong, and how badly?
- Check cost and speed per task at realistic volumes.
- Where possible, run it in shadow mode alongside the current process before anyone relies on it.
Agree the release criteria up front: the quality level, the failure types that are acceptable and the ones that are not.
Exit question: do we have evidence that it is better than what we do today?
4. Ship
Shipping is not the finish line. It is the start of the stage in which the system has to keep earning its place.
- Adoption. The people who will use it need to understand what it does, what it doesn’t do and when to override it. A system nobody trusts saves nothing.
- Monitoring. Watch quality, cost and drift over time, because inputs change and models get updated.
- Feedback. Make it easy for users to flag wrong results, and feed those cases back into the evaluation set.
- Ownership. One named owner, responsible for the system after the project team moves on.
Exit question: is it being used, and is it still delivering the number we framed?
An example: from an annual cycle to near real time
One of the systems I led replaced a transaction verification process that used to run as a manual exercise once a year.
Framing made clear what “verified” actually meant, and which mismatches genuinely needed a person to look at them. Orchestration turned the matching into an automated workflow, with the exceptions routed to the right people along with the context they needed. Evaluation compared its results with the outcomes of the manual process before anyone relied on it. And once shipped, it ran continuously instead of once a year, so problems surfaced while they could still be acted on.
None of those steps depended on a particular model. They depended on being clear about the problem, the measure and who decides.
Why the gates matter
Each stage answers a different question: should we build this, does it work, can we trust it, and is it still working? The technology changes every few months. Those questions don’t.