AI Incidents Need a Rehearsed Response
Plan how to contain an agent, preserve useful evidence, restore service and learn from mistakes before an incident occurs.

Imagine an agent starts sending an incorrect update to a connected system. The model’s answer is only one part of the incident. You also need to know which actions completed, which are pending and what customers or colleagues may now rely on.
A response plan should answer those questions before the team is under pressure. The exercise can start small, with a synthetic scenario and the actual people who would respond.
Rehearse containment at the action boundary
Identify who can suspend the workflow, revoke its credentials or disable a specific connector. Test those controls. A pause in the user interface may not stop queued jobs or actions already accepted by another service.
Decide how essential work continues while the system is unavailable. That might mean a manual route, a restricted read-only mode or a previous version. Check that the fallback can handle realistic demand.
Preserve relevant evidence with appropriate access and retention controls: affected records, action identifiers, configuration versions and timestamps. Avoid collecting entire sensitive prompts by default when a narrower record would answer the incident question.
Make reporting easier than hiding a mistake
Staff should know where to report an unexpected result or accidental disclosure. A useful first report can be incomplete. The response team can establish scope as evidence arrives.
The NCSC’s security culture guidance emphasises accessible reporting, fair treatment and learning from incidents. People need confidence that raising an innocent mistake will help the organisation respond. Read the NCSC security culture guidance.
Avoid the lazy assumption that every failure is a careless employee. Examine permissions, interface design, training, workload and the incentives surrounding the decision.
Separate recovery from the decision to restart
Restoring technical service does not automatically establish that the workflow is ready to resume. Identify the failure mode, correct the relevant controls and rerun the cases that matter. Assign a named owner to the restart decision.
Involve security, privacy, legal and business owners as appropriate to assess notification or other obligations for the actual incident. A generic AI playbook cannot determine every applicable deadline or duty.
Run a short tabletop exercise this month: one wrong action, one exposure of information and one supplier outage. Finish with assigned improvements and a date to retest them.







