The short answer
With 70 percent of a quarterly error budget spent and five weeks left, the budget is burning only slightly faster than its sustainable rate. A senior answer treats it as a decision instrument: it asks whether one incident or steady erosion consumed it, then changes how the launch ships — canary, flag, staged rollout — rather than freezing everything.
This is the scenario we put to site reliability engineering candidates, published in full on the role page. It earns a public walkthrough because nearly every candidate knows what an error budget is, and far fewer have used one to make a decision under pressure.
Your error budget for the quarter is 70 percent consumed with five weeks to go. Product wants to ship a launch that touches the checkout path. What do you do?
First, do the arithmetic
In a 13-week quarter, eight weeks have gone — about 62 percent of the window — against 70 percent of the budget. That is burning a little faster than the rate that would land exactly on budget, roughly 1.1 times. It is a warning, not an emergency.
Candidates who hear “70 percent” and reach straight for a freeze have skipped the one calculation that tells them how worried to be.
What a senior answer contains
- Treats the error budget as a decision instrument, not a report — and says who owns the call. The budget exists to settle exactly this argument between reliability and delivery. If nobody can say who decides when it is spent, the policy is decorative.
- Separates “freeze everything” from “change what we ship and how.” A launch can still go out behind a feature flag, as a canary to a small slice of traffic, or as a staged rollout with an automatic rollback threshold. Each spends less budget than a full release and keeps the decision reversible.
- Asks whether one incident or steady erosion consumed the budget. One bad outage that is now fixed leaves a very different risk picture from a service shedding a little reliability every week. The second is the one that should slow the launch, because it will keep burning.
- Has actually had this argument with a product owner, and can describe how it landed — what was traded, what shipped, what was deferred.
Where a mid answer stops
- Proposes a blanket change freeze with no path to ship. It is the safe-sounding answer, and the one that teaches product teams to ignore the error budget next time.
- Cannot say what the SLO measures or who agreed it. An error budget is only as meaningful as the objective behind it. A candidate who has never seen an SLO negotiated is quoting a number, not owning it.
- Treats the error budget as an SRE-team metric rather than a shared contract. The point of the budget is that product and engineering agreed it in advance, which is what lets it end the argument instead of starting one.
The wider point
Error budgets fail quietly. A team can track one for years without it ever changing a decision. This scenario is built to find out whether a candidate has seen the instrument actually used, which is the difference between knowing the vocabulary of SRE and having done the job.
This is the published screen for site reliability engineers. The role page carries the scenario, the full rubric and the technologies we screen on.