SkillVaultskills Browse all 500 skills

Debugging · Version 1.7.0 · Reviewed 2026-08-02

Python Async Debugger

Diagnose event-loop stalls and task cancellation with evidence, explicit trade-offs, and a verification plan.

4 method steps 6 documented failure modes 5 diagnostic checks 7 quality gates

Diagnoses blocked event loops, forgotten awaits, cancellation leaks, task lifetime mistakes, and mixed sync/async I/O in Python.

₹99 one-time

Get this skill archive

What this skill helps you do

  • Event-loop stalls
  • Task cancellation
  • Async resource leaks

How Python Async Debugger works

You provide

Shared-state access paths, pool sizing, and symptoms

It inspects

Read-modify-write and lock ordering for event-loop stalls

It decides

A task cancellation fix using atomic or constraint enforcement

You verify

Reproduce under real concurrency and confirm one outcome

What it checks first

Python Async Debugger diagnoses blocked event loops, forgotten awaits, cancellation leaks, task lifetime mistakes, and mixed sync/async I/O in Python. Use it when the work involves Event-loop stalls, Task cancellation, Async resource leaks.

  1. Every read-modify-write on shared state and whether it is atomic, locked, or transactional.
  2. Lock acquisition order across code paths, since inconsistent ordering is the definition of a deadlock risk.
  3. Whether async work outlives the request that started it, and what cancels it.
  4. Pool sizing relative to the blocking behavior of the work, because blocking calls on a small pool serialize everything.
  5. Whether the failure reproduces under load or only in production, which indicates a timing-dependent defect.

Failure modes it recognizes

  • Lost update where two transactions read the same value and the second write silently discards the first.
  • Deadlock from two paths acquiring the same two locks in opposite order.
  • Thread-pool exhaustion where blocking I/O on the pool starves the work that would release it.
  • A cancelled request whose downstream work continues, consuming capacity and producing orphaned writes.
  • Double execution of a scheduled job when two instances both believe they hold leadership.
  • Unbounded queue growth converting backpressure into memory exhaustion.

Answers it will reject

  • Adding a sleep to fix a race, which changes the probability without removing the defect.
  • Widening a lock to make a bug disappear, trading correctness ambiguity for a throughput collapse.
  • Assuming a single instance, which becomes false at the first horizontal scale event.
  • Using a database read followed by an application-level uniqueness check instead of a unique constraint.

Decision rules it applies

  • Prefer atomic database operations and unique constraints over application-level coordination.
  • Acquire locks in a globally consistent order and hold them for the shortest possible span.
  • Make operations idempotent so that at-least-once execution is safe.
  • Bound every queue and every retry; unbounded means the failure mode is chosen by the environment.

Evidence it asks for

  • Reproduce under concurrency with a load test rather than reasoning about interleavings alone.
  • Capture a thread or task dump during the stall to see what each worker is blocked on.
  • Add a unique constraint temporarily to prove whether duplicates are occurring.

The method inside

  1. Reconstruct the symptom timeline and define what healthy behavior would look like for event-loop stalls.
  2. Rank hypotheses for task cancellation by evidence, blast radius, and ability to explain every observed symptom.
  3. Run the cheapest discriminating check for async resource leaks; update confidence only when evidence changes.
  4. Separate immediate stabilization, confirmed cause, contributing conditions, and prevention; finish with a reproducible verification.

Deliverables

  • Event-loop stalls assessment
  • Task cancellation decision and action plan
  • Async resource leaks verification checklist

Evidence requirements

  • Exact symptoms and timestamps
  • Reproduction conditions and recent changes
  • Logs, traces, metrics, code, or configuration

Quality gates

  • Every material claim traces to supplied evidence or is labeled as a hypothesis.
  • The response follows the declared deliverable contract.
  • No execution, access, measurement, or verification is invented.
  • Secrets and personal data are redacted rather than repeated.
  • The user receives a concrete independent verification step.
  • The relevant failure modes in this domain were considered rather than only the reported symptom.
  • No listed anti-pattern was recommended as a solution.

Example task

Input

Our FastAPI endpoint sometimes blocks every request for several seconds even though it is declared async.

Expected output

The handler calls a synchronous database client directly on the event-loop thread. During that query no coroutine can progress; use the async driver or explicitly offload the blocking call to a bounded worker pool...

Boundaries and compatibility

Ideal for

  • Event-loop stalls: produce a decision or artifact grounded in supplied evidence.
  • Task cancellation: produce a decision or artifact grounded in supplied evidence.
  • Async resource leaks: produce a decision or artifact grounded in supplied evidence.

Out of scope

  • Guessing a root cause from a symptom alone
  • Claiming a fix worked without test evidence

Agent compatibility

  • GitHub Copilot custom agents
  • Claude Agent Skills / SKILL.md
  • Any instruction-following chat model

Tool policy: Advisory by default. No tools are assumed. If the host provides tools, use read-only evidence gathering unless the user explicitly approves a scoped write or execution action.