SkillVaultskills Browse all 500 skills

Infrastructure · Version 1.0.0 · Reviewed 2026-08-02

Kubernetes Autoscaling Designer

Review and harden metric selection and scale-up latency with evidence, explicit trade-offs, and a verification plan.

4 method steps 4 documented failure modes 4 diagnostic checks 7 quality gates

Designs horizontal, vertical, and cluster autoscaling around the metric that actually predicts saturation.

₹99 one-time

Get this skill archive

What this skill helps you do

  • Metric selection
  • Scale-up latency
  • Cluster autoscaler tuning

How Kubernetes Autoscaling Designer works

You provide

Manifests, plans, and current runtime topology

It inspects

Reversibility and blast radius for metric selection

It decides

A scale-up latency change staged by risk

You verify

Platform-native health check after each stage

What it checks first

Kubernetes Autoscaling Designer designs horizontal, vertical, and cluster autoscaling around the metric that actually predicts saturation. Use it when the work involves Metric selection, Scale-up latency, Cluster autoscaler tuning.

  1. Whether a change is reversible, and specifically whether it replaces or mutates a stateful resource.
  2. Blast radius: the number of environments, regions, and workloads a change touches at once.
  3. Identity and permission scope of the executing principal.
  4. Drift between declared and actual state.

Failure modes it recognizes

  • An immutable-attribute change forcing replacement of a stateful resource.
  • A change applied to all environments simultaneously with no canary.
  • Over-broad permissions granted to make a deployment succeed and never narrowed.
  • Manual changes creating drift that the next apply silently reverts.

Answers it will reject

  • Approving a plan from summary counts rather than reading the replacement lines.
  • Suppressing drift detection to silence noise, which disables reconciliation.
  • Granting administrative rights as a debugging shortcut.

Decision rules it applies

  • Any stateful replacement requires a tested backup and restore path before approval.
  • Roll out by blast radius: one non-critical target, then one zone, then the fleet.
  • Grant the narrowest permission that completes the task, with an expiry.

Evidence it asks for

  • Diff the plan in machine-readable form and classify every action.
  • Verify the rollback path by executing it in a non-production environment.
  • Confirm post-change health with a platform-native check, not an assumption.

The method inside

  1. Turn metric selection into explicit functional requirements and quality-attribute constraints.
  2. Model the critical path, state, trust, and failure boundaries that govern scale-up latency.
  3. Compare viable designs for cluster autoscaler tuning against weighted constraints and operational ownership.
  4. Select a design with consequences, rollout stages, observability, and a reversible adoption path.

Deliverables

  • Metric selection assessment
  • Scale-up latency decision and action plan
  • Cluster autoscaler tuning verification checklist

Evidence requirements

  • Infrastructure code or configuration
  • Runtime topology and environment constraints
  • Plan, events, policies, and failure symptoms

Quality gates

  • Every material claim traces to supplied evidence or is labeled as a hypothesis.
  • The response follows the declared deliverable contract.
  • No execution, access, measurement, or verification is invented.
  • Secrets and personal data are redacted rather than repeated.
  • The user receives a concrete independent verification step.
  • The relevant failure modes in this domain were considered rather than only the reported symptom.
  • No listed anti-pattern was recommended as a solution.

Example task

Input

Our HPA scales on CPU but latency degrades long before CPU rises, so we are always late to scale.

Expected output

CPU is the wrong signal for a workload bound by connections or downstream wait. Scale on a saturation metric such as queue depth or in-flight requests per pod, and account for the pod startup time in the target so capacity arrives before the queue builds...

Boundaries and compatibility

Ideal for

  • Metric selection: produce a decision or artifact grounded in supplied evidence.
  • Scale-up latency: produce a decision or artifact grounded in supplied evidence.
  • Cluster autoscaler tuning: produce a decision or artifact grounded in supplied evidence.

Out of scope

  • Applying infrastructure changes without approval
  • Assuming cloud access or live resource visibility

Agent compatibility

  • GitHub Copilot custom agents
  • Claude Agent Skills / SKILL.md
  • Any instruction-following chat model

Tool policy: Advisory by default. No tools are assumed. If the host provides tools, use read-only evidence gathering unless the user explicitly approves a scoped write or execution action.