nerdexam
Google

PROFESSIONAL-CLOUD-DEVOPS-ENGINEER · Question #141

Your product is currently deployed in three Google Cloud Platform (GCP) zones with your users divided between the zones. You can fail over from one zone to another, but it causes a 10-minute service d

The correct answer is B. MTTD: 5 MTTR: 20 MTBF: 90 Impact: 33%. For the new chat feature's database risk, the Mean Time to Detect (MTTD) remains 5 minutes, the Mean Time to Repair (MTTR) doubles to 20 minutes due to slower failover, the Mean Time Between Failure (MTBF) is 90 days (once per quarter), and User Impact Percentage is 33% as users

Submitted by andreas_gr· Apr 18, 2026Applying site reliability engineering principles to a service

Question

Your product is currently deployed in three Google Cloud Platform (GCP) zones with your users divided between the zones. You can fail over from one zone to another, but it causes a 10-minute service disruption for the affected users. You typically experience a database failure once per quarter and can detect it within five minutes. You are cataloging the reliability risks of a new real- time chat feature for your product. You catalog the following information for each risk: * Mean Time to Detect (MTTD) in minutes * Mean Time to Repair (MTTR) in minutes * Mean Time Between Failure (MTBF) in days * User Impact Percentage The chat feature requires a new database system that takes twice as long to successfully fail over between zones. You want to account for the risk of the new database failing in one zone. What would be the values for the risk of database failover with the new system?

Options

  • AMTTD: 5 MTTR: 10 MTBF: 90 Impact: 33%
  • BMTTD: 5 MTTR: 20 MTBF: 90 Impact: 33%
  • CMTTD: 5 MTTR: 10 MTBF: 90 Impact: 50%
  • DMTTD: 5 MTTR: 20 MTBF: 90 Impact: 50%

How the community answered

(29 responses)
  • A
    3% (1)
  • B
    76% (22)
  • C
    7% (2)
  • D
    14% (4)

Why each option

For the new chat feature's database risk, the Mean Time to Detect (MTTD) remains 5 minutes, the Mean Time to Repair (MTTR) doubles to 20 minutes due to slower failover, the Mean Time Between Failure (MTBF) is 90 days (once per quarter), and User Impact Percentage is 33% as users are divided across three zones.

AMTTD: 5 MTTR: 10 MTBF: 90 Impact: 33%

The MTTR is incorrect; it should be 20 minutes (twice the original 10 minutes for failover).

BMTTD: 5 MTTR: 20 MTBF: 90 Impact: 33%Correct

The MTTD remains 5 minutes as detection time is unchanged. The MTTR doubles from 10 to 20 minutes because the new database system takes twice as long to fail over. The MTBF is 90 days based on a quarterly failure rate, and the User Impact Percentage is 33% because users are evenly divided across three zones and one zone fails.

CMTTD: 5 MTTR: 10 MTBF: 90 Impact: 50%

The MTTR is incorrect (should be 20 minutes), and the User Impact Percentage is incorrect (should be 33% for one of three zones).

DMTTD: 5 MTTR: 20 MTBF: 90 Impact: 50%

The User Impact Percentage is incorrect; if users are divided among three zones and one fails, only 33% are impacted, not 50%.

Concept tested: Reliability metrics (MTTD, MTTR, MTBF, User Impact)

Source: https://cloud.google.com/blog/products/operations/sre-fundamentals-on-google-cloud-reliability-metrics

Topics

#SRE metrics#Reliability risk assessment#Failover#Service disruption

Community Discussion

No community discussion yet for this question.

Full PROFESSIONAL-CLOUD-DEVOPS-ENGINEER Practice