Skip to main content

Incident management

The organised response to when something goes wrong with an AI system, detecting the problem, fixing it, reporting it if required, and learning how to prevent it next time.

The Simple Version

The organised response to when something goes wrong with an AI system, detecting the problem, fixing it, reporting it if required, and learning how to prevent it next time.

Detailed Explanation

AI incident management adapts IT service management (ITIL) incident processes to the characteristics of AI failures, including non-determinism, emergent behaviour, and dual regulatory reporting obligations (EU AI Act serious incidents and GDPR personal data breaches). An AI incident management framework defines: incident triggers and detection mechanisms, triage and severity classification, investigation procedures, containment and remediation actions, regulatory notification processes, and post-incident review. The EU AI Act imposes reporting timelines for serious incidents, immediate notification for incidents causing death or critical infrastructure disruption, within 15 days for other serious harm.

Key Characteristics

  • Adapts ITIL processes to AI-specific failure modes and regulatory obligations
  • Includes severity classification aligned with EU AI Act serious-incident definition
  • Must integrate with EU AI Act Article 73 notification timelines
  • Post-incident review should feed back into risk register and risk management

Why It Matters

Organisations deploying high-risk AI should build and test incident response procedures before deployment, not after, including regulatory notification workflows and communication plans for affected individuals.

Real-World Analogy

Like a hospital's clinical incident management process, structured triage, investigation, duty of candour, and learning review that turns adverse events into safety improvements.

Common Misconceptions

  • AI incident management is just IT incident management applied to AI. AI incidents require additional steps including model investigation, bias analysis, and regulatory notification that standard IT processes do not address.
  • Only production incidents require incident management, pre-production issues identified in testing that reveal systematic safety or bias risks should also be managed and documented.

Related Terms

Sources & Further Reading