nerdexam
Google

PROFESSIONAL-CLOUD-DEVOPS-ENGINEER · Question #138

You encountered a major service outage that affected all users of the service for multiple hours. After several hours of incident management, the service returned to normal, and user access was restor

The correct answer is B. Develop a post-mortem to be distributed to stakeholders.. SRE best practice after a major incident is to write a blameless post-mortem before communicating with stakeholders. The post-mortem documents what happened, the timeline, impact, root cause, and action items. Distributing this document ensures consistent, accurate information re

Submitted by femi9· Apr 18, 2026Managing a service incident

Question

You encountered a major service outage that affected all users of the service for multiple hours. After several hours of incident management, the service returned to normal, and user access was restored. You need to provide an incident summary to relevant stakeholders following the Site Reliability Engineering recommended practices. What should you do first?

Options

  • ACall individual stakeholders to explain what happened.
  • BDevelop a post-mortem to be distributed to stakeholders.
  • CSend the Incident State Document to all the stakeholders.
  • DRequire the engineer responsible to write an apology email to all stakeholders.

How the community answered

(30 responses)
  • A
    7% (2)
  • B
    80% (24)
  • C
    3% (1)
  • D
    10% (3)

Explanation

SRE best practice after a major incident is to write a blameless post-mortem before communicating with stakeholders. The post-mortem documents what happened, the timeline, impact, root cause, and action items. Distributing this document ensures consistent, accurate information reaches all stakeholders simultaneously. Option A (individual calls) is ad-hoc, inconsistent, and not scalable for a major outage. Option C (Incident State Document) is an operational document used during the incident, not the post-incident summary. Option D (apology email) is punitive, violates blameless culture, and is not an SRE-recommended practice.

Topics

#Incident Management#Post-mortem#SRE Principles#Stakeholder Communication

Community Discussion

No community discussion yet for this question.

Full PROFESSIONAL-CLOUD-DEVOPS-ENGINEER Practice