Skip to main content
Lornets

Evidence note

An Empirical Study of Production Incidents in Generative AI Cloud Services

Peer-Reviewed Research2025Microsoft, ISSRE 2025

What was examined?

A multi-year analysis of high-severity production incidents in Microsoft GenAI cloud services, with comparison against conventional cloud incidents.

Key findings

  1. 0149.8% of the studied GenAI incidents were primarily classified as performance degradation.
  2. 0235.7% were deployment failures.
  3. 0314.5% were primarily invalid inference.
  4. 04Root causes included infrastructure, configuration, code, external usage and operational error.
  5. 0538.3% of GenAI incidents were reported by humans rather than automated monitors, compared with 13.7% in the study's non-GenAI comparison group.

Why it matters. Lornets interpretation.

AI production assurance requires both conventional system observability and AI behavioural observability. Model behaviour is only one part of the operating system.

This is the Lornets reading of the source, not a finding of the source itself.

What it does not establish

  1. 01The data comes from Microsoft GenAI cloud services.
  2. 02The exact incident counts are not published.
  3. 03Percentages should not be generalised to AI products generally.
  4. 04Microsoft's mature operational environment may differ significantly from smaller organisations.

Source

Organisation
Microsoft
Venue
ISSRE 2025
Evidence type
Peer-Reviewed Research
Published
2025
Status
Current

Read the original research

Relevant Lornets framework areas

Framework domains

  • Reliability & Recoverability
  • Observability & Operations
  • Delivery & Change Control
  • AI Assurance

Related evidence

Source record

Published
2025
Last verified
2026-08-11
Source status
Current