Debugging & Performance

Debugging Async Code and Event Loops

Async bugs rarely raise a clean traceback at the point of failure. A coroutine silently never runs, a blocking call freezes the whole loop, a task dies with its exception swallowed, or teardown trips RuntimeError: Event loop is closed long after the real mistake. The reason is structural: in asyncio, scheduling is decoupled from execution, exceptions live on Task objects until someone retrieves them, and one event loop multiplexes everything. This guide turns those invisible failures into observable ones — debug mode, slow-callback detection, never-awaited warnings, task introspection, and stepping through coroutines with pdb.

Prerequisites

  • Python 3.8+ (asyncio.all_tasks, asyncio.current_task, asyncio.run are the modern, loop-agnostic APIs; the pre-3.7 get_event_loop patterns are deprecated).
  • Python 3.9+ for asyncio.to_thread, the ergonomic way to offload a blocking call without touching an executor.
  • Python 3.11+ for the async-aware REPL (python -m asyncio) that lets you await at the prompt, and for asyncio.TaskGroup structured concurrency that re-raises the first child failure.
  • aiomonitor >= 0.7 for attaching to a live loop over a console (pip install aiomonitor); aiodebug for slow-callback logging hooks.
  • tracemalloc (stdlib) for allocation tracebacks on never-awaited coroutines.
  • Comfort reading tracebacks and driving pdb; async debugging is ordinary debugging plus scheduling, not a separate discipline.

Core concept

The event loop runs one callback at a time. A coroutine becomes a Task only when it is awaited or scheduled (create_task, gather, ensure_future); a bare some_coro() call just builds a coroutine object that does nothing until awaited — forget the await and it is garbage-collected unrun, producing the "coroutine was never awaited" warning. Because the loop is single-threaded, any synchronous blocking call (a requests GET, time.sleep, a CPU-bound loop) freezes every task, which debug mode surfaces as a slow-callback warning. Exceptions raised inside a task do not propagate to the caller; they are stored on the task and only re-raised when its result is awaited, which is why a crashed background task can vanish without a trace.

asyncio debug mode is the single most valuable switch. Enable it via PYTHONASYNCIODEBUG=1, asyncio.run(main(), debug=True), or loop.set_debug(True). It logs callbacks slower than loop.slow_callback_duration (default 0.1s), warns on never-awaited coroutines with their origin, and checks that loop calls happen on the owning thread.

Exception propagation is the second thing to internalise, because it explains most "the code just stopped" reports. When you await a coroutine directly, an exception it raises unwinds normally into the caller. When you hand a coroutine to create_task, the resulting Task runs independently: any exception is captured and parked on the task, retrieved only when something calls await task, task.result(), or task.exception(). Drop the reference to that task and let it be garbage-collected with an unretrieved exception, and the loop's default exception handler prints a "Task exception was never retrieved" message — the only breadcrumb you get. asyncio.gather(*coros) re-raises the first exception by default but leaves the siblings running unless you also cancel them; asyncio.gather(*coros, return_exceptions=True) swallows every failure into the result list instead, which is a frequent source of "why is this exception invisible" confusion. On Python 3.11+, asyncio.TaskGroup is the structured-concurrency answer: it cancels the remaining children on the first failure and raises an ExceptionGroup, so a background crash can no longer disappear.

Task lifecycle on the event loop A coroutine becomes a scheduled task, runs and suspends on awaits, then finishes or stores an exception; forgetting to await leaves it unrun. Where async tasks go wrong coroutine object created by some_coro() scheduled Task await / create_task running on loop one task at a time done: result awaited and returned exception stored silent until awaited never awaited GC'd, never ran Debug mode flags the red paths unawaited coroutines, slow callbacks, lost errors
A coroutine only runs once awaited or scheduled as a task; an exception sits on the task until its result is retrieved, and a forgotten await leaves the coroutine garbage-collected and unrun. Debug mode makes these red paths visible.

Step-by-step implementation

1. Enable debug mode

The cleanest switch is the debug flag on asyncio.run, which sets the loop into debug mode for the whole run:

Python
import asyncio

async def main() -> None:
    await asyncio.sleep(0.01)

# debug=True: log slow callbacks, warn on unawaited coroutines, check thread safety.
asyncio.run(main(), debug=True)

For a loop you manage yourself, call loop.set_debug(True); to enable it globally without code changes, export PYTHONASYNCIODEBUG=1 before launching.

2. Catch slow callbacks (blocked loop)

Any synchronous blocking call stalls the loop. Debug mode logs it:

Python
import asyncio, time

async def handler() -> None:
    time.sleep(0.5)            # BUG: synchronous sleep blocks the whole loop

async def main() -> None:
    loop = asyncio.get_running_loop()
    loop.slow_callback_duration = 0.05    # lower threshold to catch shorter stalls
    await handler()

asyncio.run(main(), debug=True)
# WARNING:asyncio:Executing <Handle ...> took 0.500 seconds

The fix is to move blocking work off the loop with await loop.run_in_executor(None, blocking_fn) or an async-native client. aiodebug.log_slow_callbacks can route these warnings into structured logging in production.

3. Find never-awaited coroutines

A forgotten await produces a RuntimeWarning at garbage-collection time — far from the bug. Promote it to an error and enable tracemalloc so the warning carries the allocation traceback:

Python
import asyncio, tracemalloc, warnings

tracemalloc.start()                                   # capture where coroutines are created
warnings.simplefilter("error", RuntimeWarning)        # turn the warning into a raised error

async def fetch() -> int:
    return 42

async def main() -> None:
    fetch()                  # BUG: missing await -> coroutine never runs
    await asyncio.sleep(0)

asyncio.run(main(), debug=True)

Run the test suite with python -W error::RuntimeWarning -X tracemalloc to fail CI on the warning with a pinpoint traceback. This technique gets a full treatment in tracing "coroutine was never awaited" warnings.

4. Introspect running tasks

When the loop hangs, ask it what is pending. asyncio.all_tasks() returns every live task; get_coro() and get_stack() show what each one is and where it is suspended:

Python
import asyncio

async def slow() -> None:
    await asyncio.sleep(3600)

async def main() -> None:
    asyncio.create_task(slow(), name="slow-worker")
    await asyncio.sleep(0)                      # let the task start and suspend
    for task in asyncio.all_tasks():
        print(task.get_name(), task.get_coro().__qualname__)
        task.print_stack()                      # where the task is parked
    print("current:", asyncio.current_task().get_name())

asyncio.run(main())

For a live, long-running service, attach aiomonitor and inspect tasks over a console without stopping the process:

Python
import aiomonitor, asyncio

async def main() -> None:
    with aiomonitor.start_monitor(asyncio.get_running_loop()):
        await asyncio.sleep(3600)               # `telnet localhost 50101`, then `ps`/`where`

5. Surface exceptions swallowed by background tasks

A fire-and-forget create_task whose coroutine raises will fail silently: the exception sits on the task and, unless something retrieves it, only ever appears as a late "Task exception was never retrieved" log line. Three mechanisms drag those failures into the light, from cheapest to most robust.

Four fates of a background-task exception A raised exception in a fire-and-forget task is only logged at garbage collection by default, but add_done_callback, loop.set_exception_handler, or asyncio.TaskGroup each surface it earlier and more reliably. Four fates of a background-task exception create_task(worker()) the coroutine raises silent default exception parked on the Task, retrieved by nobody logged only at GC add_done_callback reads task.exception() the instant the task finishes cheapest set_exception_handler catch-all net for every unretrieved failure on the loop global safety net TaskGroup 3.11+ cancels siblings, raises an ExceptionGroup structural, most robust
By default a fire-and-forget task's exception is parked on the Task and only logged at garbage collection. A done callback, a loop-level exception handler, and (best) a TaskGroup each drag the failure into the light — cheapest to most robust, left to right.

A done callback lets you inspect the outcome the moment a task finishes, without awaiting it:

Python
import asyncio

def report(task: asyncio.Task) -> None:
    if task.cancelled():
        return
    exc = task.exception()                  # returns the stored exception, or None
    if exc is not None:
        # Re-raise, log, or route to your error tracker here.
        raise exc

async def worker() -> None:
    raise ValueError("boom")                # would otherwise vanish

async def main() -> None:
    task = asyncio.create_task(worker())
    task.add_done_callback(report)          # fires when the task completes
    await asyncio.sleep(0.1)

asyncio.run(main())                         # ValueError now surfaces via the callback

A loop-level exception handler is the catch-all safety net for every task you forgot to guard. It receives exactly the retrieval failures that debug mode would otherwise bury in the log:

Python
import asyncio

def handler(loop: asyncio.AbstractEventLoop, context: dict) -> None:
    exc = context.get("exception")
    # `context["message"]` and `context.get("task")` pinpoint the culprit.
    print("unhandled loop error:", repr(exc), context["message"])

async def main() -> None:
    asyncio.get_running_loop().set_exception_handler(handler)
    asyncio.create_task(asyncio.sleep(-1))  # ValueError: sleep length must be non-negative
    await asyncio.sleep(0.1)

asyncio.run(main())

asyncio.TaskGroup (Python 3.11+) is the structural fix: children that raise cancel their siblings, and the group raises an ExceptionGroup at the end of the async with, so nothing is ever left unretrieved.

Python
import asyncio

async def ok() -> None:
    await asyncio.sleep(0.05)

async def bad() -> None:
    raise RuntimeError("child failed")

async def main() -> None:
    async with asyncio.TaskGroup() as tg:   # 3.11+
        tg.create_task(ok())
        tg.create_task(bad())               # cancels `ok`, propagates as ExceptionGroup

asyncio.run(main())                         # raises ExceptionGroup('unhandled errors ...')

Prefer TaskGroup for new code; reach for add_done_callback or set_exception_handler when retrofitting a long-lived service you cannot restructure. The teardown counterpart — pending tasks that outlive the loop — is covered in the dedicated "Event loop is closed" guide.

6. Step through with pdb

breakpoint() works inside a coroutine; the loop pauses while you are at the prompt (which can itself trip slow-callback warnings — expected). On Python 3.11+, the async REPL (python -m asyncio) lets you await expressions at the prompt to inspect coroutine results interactively:

Python
import asyncio

async def compute(x: int) -> int:
    result = x * 2
    breakpoint()              # pdb here; `p result`, `await some_coro()` on 3.11+ REPL
    return result

asyncio.run(compute(21))

When the breakpoint must live inside an async test, scope the loop correctly first — see how to scope pytest fixtures for async tests and the pytest-asyncio vs anyio scoping trade-offs so the breakpoint runs on a live loop. General pdb mechanics live in interactive debugging with pdb and ipdb.

Verification

  • Debug mode is on: print(asyncio.get_running_loop().get_debug()) returns True inside main.
  • Slow callbacks are caught: the "Executing ... took N seconds" warning fires for a deliberately blocking call once you lower slow_callback_duration.
  • Unawaited coroutines fail loudly: running under -W error::RuntimeWarning turns a missing await into a raised error rather than a late GC warning.
  • Background failures are not swallowed: a task that raises reaches your add_done_callback, set_exception_handler, or TaskGroup boundary instead of only logging "Task exception was never retrieved" at GC time.
  • No task leaks at shutdown: after asyncio.run returns, asyncio.all_tasks() is empty (it raises outside a loop, so check inside a final await). Pending tasks at teardown are the classic cause of the next error.

Troubleshooting

SymptomRoot causeFix
RuntimeWarning: coroutine '...' was never awaitedA coroutine was created but never awaited or scheduledawait it, asyncio.create_task(...), or pass to gather; run with tracemalloc to find it
Executing <Handle ...> took N secondsSynchronous blocking call on the loopMove work to run_in_executor or an async client
Loop hangs with no errorA task is awaiting something that never resolvesasyncio.all_tasks() + task.print_stack() to find the parked task
Background task crashed silentlyException stored on the task, never retrievedAdd a done callback, install loop.set_exception_handler, or group work under asyncio.TaskGroup (step 5); await/gather also retrieves it
gather returned an exception object instead of raisingCalled with return_exceptions=True, which packs failures into the result listInspect each result with isinstance(r, Exception), or drop the flag so the first failure propagates
RuntimeError: Event loop is closed at teardownReusing or scheduling on a closed loopSee the dedicated guide below
RuntimeError: no running event loopCalling create_task/get_running_loop outside a coroutineCall inside an async def driven by asyncio.run

Seeing what the loop is actually doing

Async bugs are hard less because the concurrency is complicated than because the loop is opaque: nothing prints, nothing blocks, and the symptom is a request that never completes. Three built-in observability tools remove most of that opacity, and none of them requires a debugger.

Debug mode. asyncio.run(main(), debug=True) (or PYTHONASYNCIODEBUG=1) enables four behaviours at once: coroutines that are never awaited are reported with their creation traceback, callbacks that take longer than 100ms are logged as slow, tasks destroyed while pending raise a warning, and the loop checks that non-threadsafe calls come from the right thread. The slow-callback log alone finds most accidental blocking calls, because a synchronous database driver inside a coroutine shows up as a callback that took 400ms.

Task inspection. asyncio.all_tasks() returns every live task, and each task's get_coro() and get_stack() show where it is suspended. Printing that set on a timer turns a hung service into a readable list of what is waiting on what.

Python
import asyncio

async def dump_tasks(interval: float = 5.0):
    while True:
        await asyncio.sleep(interval)
        for task in asyncio.all_tasks():
            if task.done():
                continue
            frame = task.get_stack(limit=1)
            where = frame[0].f_code.co_name if frame else "<no frame>"
            print(f"{task.get_name():<20} suspended in {where}")

Task names. The single cheapest change to async debuggability is naming every task at creation: asyncio.create_task(poll(), name=f"poll:{shard}"). Unnamed tasks appear as Task-17, which tells you nothing; named ones make both the dump above and every warning message self-explanatory.

Two structural habits reduce the number of times you need any of this. Prefer asyncio.TaskGroup (3.11+) or anyio's task groups over bare create_task, because a task group propagates exceptions and cancels siblings rather than letting a failed background task disappear silently. And hold a reference to every task you do create — the event loop keeps only a weak reference, so a task whose only reference is a local variable can be garbage-collected mid-flight, producing the notorious "Task was destroyed but it is pending" with no obvious cause.

Four observability levers for a running loop A stack of four asyncio observability tools: debug mode, the all_tasks inspection API, task naming, and task groups, each with what it reveals and what it costs. Four observability levers for a running loop debug mode unawaited coroutines, slow callbacks, wrong-thread calls all_tasks() dump what every live task is suspended on named tasks turns Task-17 into poll:shard-3 in every message task groups exceptions propagate instead of vanishing Debug mode is safe to enable in the whole test suite and costs only runtime.
The first three are diagnostics; the fourth is the structural change that stops the bug being possible.

A related habit for services: log the task name in every request-scoped log line. With named tasks and a contextvar carrying the request id, a hung request can be traced from its log lines to the exact suspended frame in the task dump, which turns a class of production incident that normally requires a reproduction into a five-minute read.

Finally, be deliberate about where cancellation is handled. A coroutine that catches bare Exception will swallow CancelledError on Python 3.7 and earlier, and on 3.8+ it inherits from BaseException precisely so that a broad except Exception no longer eats it. Code that must clean up on cancellation should catch asyncio.CancelledError explicitly and re-raise after cleaning up, because a swallowed cancellation turns a task-group shutdown into a hang that looks identical to a deadlock.

Frequently Asked Questions

How do I enable asyncio debug mode? Set the environment variable PYTHONASYNCIODEBUG=1, pass debug=True to asyncio.run, or call loop.set_debug(True) on a running loop. Debug mode logs slow callbacks, warns about coroutines that were never awaited, and checks that calls happen on the right thread.

Why do I get a coroutine was never awaited warning? You called an async function but never awaited the coroutine it returned or scheduled it as a task, so it was garbage collected without running. Await it, wrap it in asyncio.create_task, or pass it to asyncio.gather.

How do I see every task currently running on the event loop? Call asyncio.all_tasks(loop) to get the set of pending tasks, and asyncio.current_task() for the one running now. Each task's get_coro and get_stack methods reveal what it is and where it is suspended.

Can I use pdb inside an async function? Yes. breakpoint() works inside a coroutine on Python 3.7+, and on 3.11+ the asyncio REPL plus pdb handles awaits at the prompt. The loop is paused while you are at the breakpoint, so long pauses can trip slow-callback warnings.

Why did my background task fail without any error? An exception raised inside a fire-and-forget task is stored on the Task and only re-raised when its result is awaited. If nothing awaits it and it is garbage collected, asyncio emits a "Task exception was never retrieved" message via the loop exception handler. Await or gather the task, attach add_done_callback to inspect task.exception(), install loop.set_exception_handler, or group work under asyncio.TaskGroup so the first failure propagates.

← Back to Systematic Debugging & Performance Profiling