The harness must evolve with the model and the business
Keep interfaces stable, measure changes and retire controls that no longer help.

A harness is never finished simply because the first model integration works. The model changes, the business changes, APIs change and users discover cases absent from the original brief. I want the client to inherit a system that can improve through a controlled process.
Identify behaviour and compare changes
First, make current behaviour identifiable. Record the model and configuration, prompt version, tool contracts, relevant data sources and release version. A report that “the AI feels different” becomes easier to investigate when the team can connect it to a specific change.
Second, preserve a meaningful comparison. Before adopting a new model or changing the workflow, run representative cases against existing and proposed versions. Compare quality, cost, latency, escalation and action correctness. Keep difficult cases visible; averages can conceal the failure that matters to an operator.
Model improvements can make an old workaround unnecessary. They can also open a capability the architecture cannot yet use. Anthropic’s managed-agents work discusses assumptions in harnesses becoming stale, while its application-development research recommends re-examining the harness as models improve. Anthropic — Scaling Managed Agents Anthropic — Harness design for long-running application development
Evolve the parts that earn their place
My interpretation is that addition and removal should both require evidence. Do not preserve a complicated routing layer merely because it was expensive to build. Do not remove a permission boundary because the latest model appears more trustworthy. A workaround for model behaviour and a business control serve different purposes.
I would separate stable interfaces from components expected to evolve. Business rules, approved actions and the definition of a completed task need clear owners. Model adapters, retrieval settings and implementation strategies can change behind those contracts when evaluation supports the change.
Give improvement an owner and a release path
Continuous improvement also needs a release path. Propose a change, test it, review the evidence, deploy to a limited scope where appropriate, and monitor real outcomes. Keep a route back to the known version. “The agent improves itself” is not enough information for a person responsible for the service.
The operating team needs useful signals rather than an ocean of logs. Which cases consume the most review? Where does retrieval fail? Which tool is unavailable? Which fallback is used? Store necessary evidence with appropriate access and retention, especially when prompts or traces can contain personal or confidential information.
Finally, agree who owns the harness after launch. Someone must maintain tools, review failures, update evaluation cases and decide when a new version is ready. This is part of the product, not administration to add later.
The long-term value of embedded AI delivery is a working capability the organisation can understand, measure and evolve. A good handover lets the team keep improving the system without starting the conversation from zero every time the model changes.
Embedded AI Delivery includes the architecture, operating model and handover needed beyond the first release. fdo.codes







