SkillVaultskills Browse all 500 skills

Reliability · Version 1.4.0 · Reviewed 2026-08-02

Operational Runbook Author

Reduce production risk in runbook structure and decision point design with evidence, explicit trade-offs, and a verification plan.

4 method steps 6 documented failure modes 5 diagnostic checks 7 quality gates

Writes runbooks that work at 3am: explicit preconditions, decision points, verification, and a defined escalation boundary.

₹149 one-time

Get this skill archive

What this skill helps you do

  • Runbook structure
  • Decision point design
  • Escalation boundary

How Operational Runbook Author works

You provide

Incident notes, alert definitions, and prior variants

It inspects

Where a responder must branch and what discriminates

It decides

Ordered steps with an explicit escalation boundary

You verify

A responder unfamiliar with the system completes it

What it checks first

Operational Runbook Author writes runbooks that work at 3am: explicit preconditions, decision points, verification, and a defined escalation boundary. Use it when the work involves Runbook structure, Decision point design, Escalation boundary.

  1. Where the token is validated and whether the signature, issuer, audience, and expiry are all checked.
  2. Token lifetime relative to the revocation requirement, since a long-lived token cannot be revoked without a check.
  3. Whether authorization is evaluated at the data layer or only at the route.
  4. Token storage location, which determines exposure to XSS versus CSRF.
  5. Whether the flow uses PKCE and validates `state`, which decides interception and CSRF resistance.

Failure modes it recognizes

  • Accepting a JWT with `alg: none` or verifying with the wrong key type, allowing forged tokens.
  • Validating signature but not `aud`, allowing a token minted for another service to be replayed.
  • Refresh tokens without rotation or reuse detection, so a stolen token remains valid indefinitely.
  • Session fixation from not regenerating the session identifier after privilege change.
  • Role checks in the UI only, with the API trusting a client-supplied role claim.
  • Redirect URI matched by prefix rather than exact value, enabling token exfiltration.

Answers it will reject

  • Building a custom authentication protocol instead of using a reviewed standard implementation.
  • Storing access tokens in `localStorage` while claiming XSS is out of scope.
  • Using long-lived API keys where short-lived scoped credentials are available.
  • Treating authentication as sufficient for authorization, which is the most common access-control defect.

Decision rules it applies

  • Authorize at the point of data access, because every route eventually gets a second caller.
  • Prefer short-lived access tokens with rotating refresh tokens and reuse detection.
  • Match redirect URIs exactly and require PKCE for every public client.
  • Fail closed: an unparseable or unverifiable token is a rejection, never a fallback to anonymous.

Evidence it asks for

  • Decode a real token and verify each claim check exists in code: signature, issuer, audience, expiry, and not-before.
  • Test horizontal access by requesting another tenant's resource with a valid token.
  • Confirm revocation actually terminates access within the documented window.

The method inside

  1. Turn runbook structure into explicit functional requirements and quality-attribute constraints.
  2. Model the critical path, state, trust, and failure boundaries that govern decision point design.
  3. Compare viable designs for escalation boundary against weighted constraints and operational ownership.
  4. Select a design with consequences, rollout stages, observability, and a reversible adoption path.

Deliverables

  • Runbook structure assessment
  • Decision point design decision and action plan
  • Escalation boundary verification checklist

Evidence requirements

  • User-visible symptoms and SLO impact
  • Timeline, telemetry, deploys, and dependency state
  • Current mitigations and operational constraints

Quality gates

  • Every material claim traces to supplied evidence or is labeled as a hypothesis.
  • The response follows the declared deliverable contract.
  • No execution, access, measurement, or verification is invented.
  • Secrets and personal data are redacted rather than repeated.
  • The user receives a concrete independent verification step.
  • The relevant failure modes in this domain were considered rather than only the reported symptom.
  • No listed anti-pattern was recommended as a solution.

Example task

Input

Turn our incident notes into a runbook for the on-call rotation.

Expected output

Notes describe what happened once; a runbook must handle the variants. Add the discriminating check at each branch so a responder can tell which case they are in, and state the escalation point explicitly, since the most expensive runbook failure is someone continuing past their confidence...

Boundaries and compatibility

Ideal for

  • Runbook structure: produce a decision or artifact grounded in supplied evidence.
  • Decision point design: produce a decision or artifact grounded in supplied evidence.
  • Escalation boundary: produce a decision or artifact grounded in supplied evidence.

Out of scope

  • Replacing incident command authority
  • Calling a trigger the root cause without a causal chain

Agent compatibility

  • GitHub Copilot custom agents
  • Claude Agent Skills / SKILL.md
  • Any instruction-following chat model

Tool policy: Advisory by default. No tools are assumed. If the host provides tools, use read-only evidence gathering unless the user explicitly approves a scoped write or execution action.