Detect
Capture the output, triggering input, workflow, timestamp, model and prompt versions, affected user or system, and the signal that raised concern. Avoid rewriting the evidence before preservation.
Human incident command
A practical operating model for detecting risky AI outputs, containing impact, assigning accountable reviewers, preserving evidence, and closing incidents with explicit decisions.
Severity and response contract
| Level | Criteria | Immediate action | Acknowledge |
|---|---|---|---|
| SEV-1 | Active or imminent material harm, data exposure, unauthorized action, or broad customer impact. | Contain immediately | 15 min |
| SEV-2 | High-risk output reached a limited audience or a control failed with credible impact. | Pause affected lane | 30 min |
| SEV-3 | Policy breach or quality failure with bounded impact and no evidence of ongoing harm. | Queue human review | 4 hours |
| SEV-4 | Low-risk anomaly, near miss, or monitoring signal requiring trend review. | Record and monitor | 2 business days |
Incident lifecycle
Capture the output, triggering input, workflow, timestamp, model and prompt versions, affected user or system, and the signal that raised concern. Avoid rewriting the evidence before preservation.
Stop or narrow the affected workflow, revoke unsafe tool access, suppress the specific output, and protect users or data. Containment should be reversible and proportional to the observed risk.
Name one incident owner and the required specialists. Define response deadlines, decision authority, communication responsibility, and the conditions that require legal, security, clinical, or executive escalation.
Reconstruct the event from immutable evidence. Identify the failed control, similar exposures, blast radius, contributing configuration changes, and whether the issue is reproducible.
Record remediation, residual risk, customer communication, release constraints, and the human decision to resolve, monitor, or escalate. Do not close an incident because the alert stopped firing.
Convert the failure into a policy update, test case, monitoring rule, ownership change, or product control. Track the corrective action separately from the immediate incident response.
Every incident needs one accountable owner, one current status, one next deadline, and one person authorized to accept residual risk. Contributors do not replace ownership.
Preserve impact, root cause, control changes, validation results, communications, accepted residual risk, and follow-up work. Resolution is a decision, not a deleted alert.