Frame, orchestrate, evaluate, ship: how I take AI from idea to production
The four stages I use to turn an AI idea into a system people rely on, and the question each stage has to answer before the next one starts.
Most AI projects that stall don’t fail because of the model. They fail at the edges: a problem that was never sharply defined, no way to tell whether the system is better than what it replaces, or nobody owning it after launch.
To avoid those failures, I take every AI initiative through four stages: frame, orchestrate, evaluate and ship. Each stage ends with a question that has to be answered before the next one starts. Skipping a stage doesn’t save time; it just moves the cost to later, when it is more expensive.
1. Frame
Start with the business problem, not the technology. The questions at this stage:
- What should change, and how will we measure it? Fewer hours spent, faster turnaround, fewer errors, earlier detection. Pick one primary number and record today’s baseline.
- Where is human judgement required? Some decisions should stay with people for reasons of accountability, regulation or trust. Name them now, not after the design is done.
- What data exists, and can we use it? Availability, quality, ownership and data protection rules all shape what is possible.
- What are the risks? Wrong answers, biased outcomes, security, and obligations under rules such as the GDPR and the EU AI Act.
The output is a short brief: the problem, the metric, the baseline, the boundaries and the owner. If that brief can’t be written, the project isn’t ready for a design.
Exit question: do we agree on what success looks like, and how we’ll know?
2. Orchestrate
Now design how it works. This is where the architecture decisions are made:
- Workflow or agent? If the steps are known in advance, a workflow is simpler and more reliable. Agents are for the parts where the path depends on what is found along the way.
- Tools and data access. Which systems does it read from and act on, with which permissions? Expose them as narrow, well-described tools.
- Guardrails. What must it never do, which actions need approval, and what happens when it is unsure?
- Human checkpoints. Where people review, approve or take over, and who those people are.
AI agents are part of how I build these systems too. Coding agents turn the design into working software quickly, while I direct the process, set the architecture and review what they produce.
The output is a working prototype running on real data, not a slide deck.
Exit question: does it work on real cases, end to end?
3. Evaluate
This is the stage most often skipped, and the one that decides whether a system can be trusted.
- Build an evaluation set from real cases with known correct outcomes, including the difficult ones.
- Measure against the baseline from the framing stage, not against an ideal.
- Study the failures, not just the score. Which kinds of cases go wrong, and how badly?
- Check cost and speed per task at realistic volumes.
- Where possible, run it in shadow mode alongside the current process before anyone relies on it.
Agree the release criteria up front: the quality level, the failure types that are acceptable and the ones that are not.
Exit question: do we have evidence that it is better than what we do today?
4. Ship
Shipping is not the finish line. It is the start of the stage in which the system has to keep earning its place.
- Adoption. The people who will use it need to understand what it does, what it doesn’t do and when to override it. A system nobody trusts saves nothing.
- Monitoring. Watch quality, cost and drift over time, because inputs change and models get updated.
- Feedback. Make it easy for users to flag wrong results, and feed those cases back into the evaluation set.
- Ownership. One named owner, responsible for the system after the project team moves on.
Exit question: is it being used, and is it still delivering the number we framed?
An example: from an annual cycle to near real time
One of the systems I led replaced a transaction verification process that used to run as a manual exercise once a year.
Framing made clear what “verified” actually meant, and which mismatches genuinely needed a person to look at them. Orchestration turned the matching into an automated workflow, with the exceptions routed to the right people along with the context they needed. Evaluation compared its results with the outcomes of the manual process before anyone relied on it. And once shipped, it ran continuously instead of once a year, so problems surfaced while they could still be acted on.
None of those steps depended on a particular model. They depended on being clear about the problem, the measure and who decides.
Why the gates matter
Each stage answers a different question: should we build this, does it work, can we trust it, and is it still working? The technology changes every few months. Those questions don’t.