SkillVaultskills Browse all 500 skills

Reliability · Version 1.2.0 · Reviewed 2026-08-02

SLO & Error Budget Coach

Reduce production risk in SLI definition and SLO target selection with evidence, explicit trade-offs, and a verification plan.

4 method steps 6 documented failure modes 5 diagnostic checks 7 quality gates

Turns user journeys into measurable SLIs, defensible SLO targets, burn-rate alerts, and release policies tied to error-budget consumption.

₹99 one-time

Get this skill archive

What this skill helps you do

  • SLI definition
  • SLO target selection
  • Burn-rate alerting

How SLO & Error Budget Coach works

You provide

Cost breakdown by tag, usage data, and growth trend

It inspects

Unit cost and idle capacity behind SLI definition

It decides

A SLO target selection action with a reliability guardrail

You verify

Cost per thousand requests tracked after the change

What it checks first

SLO & Error Budget Coach turns user journeys into measurable SLIs, defensible SLO targets, burn-rate alerts, and release policies tied to error-budget consumption. Use it when the work involves SLI definition, SLO target selection, Burn-rate alerting.

  1. Unit cost per business transaction rather than total spend, because total spend rises with healthy growth.
  2. The split between compute, storage, network egress, and managed-service premiums.
  3. Idle versus utilized capacity, which distinguishes a sizing problem from an architecture problem.
  4. Whether cost scales with traffic, with data retained, or with time — each has a different lever.
  5. Cross-zone and cross-region traffic, which is frequently the largest unattributed line item.

Failure modes it recognizes

  • Over-provisioned requests in a scheduler reserving capacity that is never used but is fully billed.
  • Log and metric retention growing without a policy until observability costs exceed the workload.
  • Cross-AZ chatter between services that could be zone-aligned, billed per gigabyte in both directions.
  • Orphaned resources — unattached volumes, idle load balancers, old snapshots — with no owner.
  • A development environment running production-sized infrastructure continuously.
  • Data egress from object storage to the internet where a CDN would serve the same bytes far cheaper.

Answers it will reject

  • Cutting cost by reducing redundancy, which trades a predictable bill for an unpredictable outage.
  • Optimizing the largest line item when a smaller one has the worse unit economics and faster growth.
  • Committing to reserved capacity before the workload shape is stable.
  • Reporting savings without a reliability guardrail, so a regression is discovered by customers.

Decision rules it applies

  • Attribute cost to a team or product before optimizing; unattributed cost never gets reduced.
  • Prefer eliminating waste over renegotiating price — waste compounds and discounts do not.
  • Any cost reduction must state its reliability and performance guardrail and how a regression is detected.
  • Rightsize using p95 utilization over a full business cycle, not a single day.

Evidence it asks for

  • Break cost down by tag, namespace, and service, and report cost per thousand requests.
  • Compare requested versus used resources over a representative window.
  • Track the trend of unit cost, since a flat unit cost during growth is success.

The method inside

  1. Establish the current state and the constraint that actually limits SLI definition.
  2. Separate the requested solution from the underlying problem in SLO target selection, and name the assumptions carrying the most risk.
  3. Compare only viable options for burn-rate alerting against weighted constraints, cost of reversal, and operational ownership.
  4. Commit to a sequenced recommendation with success criteria, guardrails, and the observation that would reverse it.

Deliverables

  • SLI definition assessment
  • SLO target selection decision and action plan
  • Burn-rate alerting verification checklist

Evidence requirements

  • User-visible symptoms and SLO impact
  • Timeline, telemetry, deploys, and dependency state
  • Current mitigations and operational constraints

Quality gates

  • Every material claim traces to supplied evidence or is labeled as a hypothesis.
  • The response follows the declared deliverable contract.
  • No execution, access, measurement, or verification is invented.
  • Secrets and personal data are redacted rather than repeated.
  • The user receives a concrete independent verification step.
  • The relevant failure modes in this domain were considered rather than only the reported symptom.
  • No listed anti-pattern was recommended as a solution.

Example task

Input

Help us replace our 99.99 percent uptime goal with an SLO that reflects whether customers can successfully publish data.

Expected output

Measure valid publish attempts that complete durably within the latency threshold, excluding client validation failures. Use a rolling 28-day target informed by current performance and business tolerance...

Boundaries and compatibility

Ideal for

  • SLI definition: produce a decision or artifact grounded in supplied evidence.
  • SLO target selection: produce a decision or artifact grounded in supplied evidence.
  • Burn-rate alerting: produce a decision or artifact grounded in supplied evidence.

Out of scope

  • Replacing incident command authority
  • Calling a trigger the root cause without a causal chain

Agent compatibility

  • GitHub Copilot custom agents
  • Claude Agent Skills / SKILL.md
  • Any instruction-following chat model

Tool policy: Advisory by default. No tools are assumed. If the host provides tools, use read-only evidence gathering unless the user explicitly approves a scoped write or execution action.