Agentic AI without the hype: where agents earn their place
An agent is a language model, a set of tools and a loop. Where that combination pays off in real organisations, and where a plain workflow is still the better choice.
“Agent” has become one of the most stretched words in technology. It gets attached to chatbots, scripts with a model call inside, and genuinely autonomous systems alike. That makes it hard to decide where agents belong in an organisation, and easy to build one where something simpler would do better.
Here is the definition I work with: an agent is a language model, a set of tools and a loop. The model looks at the goal and what it knows so far, decides on the next step, calls a tool, reads the result and repeats until the job is done or it hands over to a person. Everything else (memory, planning, multiple agents talking to each other) builds on that loop.
Workflow or agent?
The first question is not which agent to build, but whether you need one at all.
If you can draw the steps in advance, build a workflow: a fixed sequence where a model handles specific steps, such as reading a document or classifying a request, and ordinary code handles the rest. Workflows are predictable, cheap to run and easy to test.
Reach for an agent when the path depends on what you find along the way. Investigating an incident, chasing down why two records don’t match or assembling context from several systems are hard to script, because the next step depends on the last result.
In practice, most production systems I build are workflows with agentic pockets: a deterministic backbone, with an agent handling the steps that genuinely need judgement.
Where agents earn their place
The strongest use cases I see share three traits: high volume, messy inputs and a clear definition of a good outcome.
- Triage and classification. Reports arrive in different formats and languages, and each one needs a category, a priority and an owner. An agent can read, enrich and route them in seconds, so people spend their time on the cases that matter.
- Investigation. Pulling context from case histories, documents and databases to prepare a decision is exactly the kind of multi-step, path-dependent work agents handle well.
- Reconciliation. Matching records across systems and explaining the exceptions turns a periodic manual exercise into something that can run continuously.
- Document review. Flagging documents that look manipulated, incomplete or inconsistent, for a person to check, rather than making the final call.
Notice that none of these remove the human. They change where the human spends their attention.
The parts that matter more than the model
When an agent misbehaves, the model is rarely the root cause. These usually are:
- Tools. Each tool should do one thing, be clearly described and only have the permissions it needs. Standards such as the Model Context Protocol (MCP) make it much easier to expose systems to agents consistently, but the design discipline is still yours.
- Context. An agent can only reason about what it can see. Decide what it retrieves, what it remembers between runs and what it should never see.
- Guardrails. Spell out what the agent may never do, cap how many steps and how much budget it can use, and require approval before anything irreversible.
- Observability. Log every step and tool call. When something goes wrong, you need to replay the run, not guess.
- Evaluation. A set of real cases with known good outcomes, run before every change. Without it, you are tuning by anecdote.
Humans in the loop, on purpose
“Human in the loop” is often added at the end as a safety blanket. It works better when it is designed in from the start, at the points where judgement or accountability actually sits:
- approval before actions that are hard to undo, such as payments, external messages or record changes
- a review queue for low-confidence results, with the agent’s reasoning attached
- a clear escalation path when the agent is stuck, instead of letting it try forever
The goal is not to remove people. It is to move them to the decisions that need them.
The same holds when the agent writes the code
AI coding agents now write much of the code on my client builds. The pattern is the same as in any agentic system: the agent handles volume, and I own the framing, the architecture, the review and the final engineering call. The agent is fast; the direction still has to come from someone who understands the problem and is accountable for the result.
A short checklist before you build one
- Can you write down what a good outcome looks like, and check it on real examples?
- Could a workflow do the job? If so, start there.
- Does each tool have the narrowest permissions that still work?
- Where are the human checkpoints, and who staffs them?
- Can you replay any run from its logs?
- Who owns the system after launch?
If you can answer all six, an agent is very likely worth building. If you can’t, the gaps are where the project would have stalled anyway.