Skip to main content

Evidence note

An Empirical Study of Production Incidents in Generative AI Cloud Services

Peer-Reviewed Research2025Microsoft, ISSRE 2025

What was examined?

A multi-year analysis of high-severity production incidents in Microsoft GenAI cloud services, with comparison against conventional cloud incidents.

Key findings

  1. 149.8% of the studied GenAI incidents were primarily classified as performance degradation.
  2. 235.7% were deployment failures.
  3. 314.5% were primarily invalid inference.
  4. 4Root causes included infrastructure, configuration, code, external usage and operational error.
  5. 538.3% of GenAI incidents were reported by humans rather than automated monitors, compared with 13.7% in the study's non-GenAI comparison group.

Why it matters. Lornets interpretation.

AI production assurance requires both conventional system observability and AI behavioural observability. Model behaviour is only one part of the operating system.

This is the Lornets reading of the source, not a finding of the source itself.

What it does not establish

  1. 1The data comes from Microsoft GenAI cloud services.
  2. 2The exact incident counts are not published.
  3. 3Percentages should not be generalised to AI products generally.
  4. 4Microsoft's mature operational environment may differ significantly from smaller organisations.

Source

Organisation
Microsoft
Venue
ISSRE 2025
Evidence type
Peer-Reviewed Research
Published
2025
Status
Current

Read the original research

Relevant Lornets framework areas

Framework domains

  • Reliability & Recoverability
  • Observability & Operations
  • Delivery & Change Control
  • AI Assurance

Related evidence

Last verified 2026-08-11

2024

Government Guidance

NIST AI 600-1: Generative Artificial Intelligence Profile

The profile addresses risks including confabulation, data privacy, information integrity, information security and component integration.

National Institute of Standards and Technology

  • AI Assurance
  • Security & Supply Chain