Asynchronous programming
Be able to write async code with asyncio and understand when concurrency pays off — as in an API that calls an LLM.
Prerequisites
- DPython — classes and objectsrequired
Intuition
An API that calls an LLM waits 2 seconds per answer — and does nothing meanwhile. With asynchronous code the function hands back control while it waits (await), so that other requests can be served. One thread, thousands of pending calls.
It helps when the bottleneck is waiting (network, disk, database). It does not help when the bottleneck is computation (matrix multiplication) — there you need processes or a GPU.
The rules: an async def has to be awaited; a blocking call (time.sleep, requests.get) inside async freezes everything — use asyncio.sleep, httpx.AsyncClient. Limit the concurrency (a semaphore) so that you do not fire off 10 000 calls at once.
Code
import asyncio, time, httpx
async def ask_llm(client, sem, prompt):
async with sem: # at most N concurrent
r = await client.post(URL, json={"prompt": prompt}, timeout=30)
r.raise_for_status()
return r.json()["text"]
async def main(prompts):
sem = asyncio.Semaphore(8)
async with httpx.AsyncClient() as client:
t = time.perf_counter()
answers = await asyncio.gather(*(ask_llm(client, sem, p) for p in prompts), return_exceptions=True)
print(f"{len(prompts)} calls in {time.perf_counter() - t:.1f} s")
return answers
asyncio.run(main(["hi"] * 40)) # ~10 s with 8 concurrent at 2 s each, instead of 80 s
return_exceptions=True means that one failure does not bring all of them down — you get the exception as a value and can handle it per prompt.
Mastery means
- Writes async functions and runs them concurrently with gather
- Explains when async pays off (I/O-bound) and when it does not (CPU-bound)
- Limits the concurrency with a semaphore and handles timeouts
Sign in to do the exercises and build your mastery up.
Sources
- Python — asyncio (PSF) — PSF
- httpx — dokumentation (BSD-3) — BSD-3-Clause