PROFESSIONAL-CLOUD-DEVOPS-ENGINEER · Question #173
You are on-call for an infrastructure service that has a large number of dependent systems. You receive an alert indicating that the service is failing to serve most of its requests and all of its dep
The correct answer is C. Establish a communication channel where incident responders and leads can communicate with. After establishing an incident command structure, the immediate next step is to set up a dedicated communication channel to facilitate real-time collaboration among the incident response team.
Question
Options
- ALook for ways to mitigate user impact and deploy the mitigations to production.
- BContact the affected service owners and update them on the status of the incident.
- CEstablish a communication channel where incident responders and leads can communicate with
- DStart a postmortem, add incident information, circulate the draft internally, and ask internal
How the community answered
(57 responses)- A5% (3)
- B4% (2)
- C79% (45)
- D12% (7)
Why each option
After establishing an incident command structure, the immediate next step is to set up a dedicated communication channel to facilitate real-time collaboration among the incident response team.
While mitigating user impact is a priority, it is typically the Operations Lead's responsibility to identify and deploy mitigations after initial assessment, and not the IC's immediate next step after role assignment.
Contacting affected service owners and providing updates is the responsibility of the Communications Lead, and this action is secondary to establishing internal communication channels for the incident team.
Once the core incident roles (IC, OL, CL) are established, setting up a dedicated communication channel (e.g., chat room, war room) is crucial. This provides a centralized place for all incident responders and leads to communicate updates, share findings, coordinate actions, and make decisions efficiently, which is foundational for effective incident resolution.
Starting a postmortem is an activity that occurs *after* the incident is resolved or significantly stabilized, not as an immediate next step during an active, high-severity incident.
Concept tested: SRE incident response process (initial steps)
Source: https://cloud.google.com/sre/docs/incidents#incident_commander
Topics
Community Discussion
No community discussion yet for this question.