Threads vs coroutines: who decides to yield

Budding

6 min read

Threads and coroutines both let one program do many things at once, so they get filed together and treated as interchangeable. They are not. The difference is who is in control - who decides when a running task steps aside - and that single choice cascades into what each task costs, how many you can have, and how they fail.

The unit, and what it costs

  • A thread is scheduled by the operating system. It gets its own stack (often a megabyte or so), and switching between threads means a trip into the kernel to save and restore state. That cost is small in isolation and ruinous in bulk: you can run thousands of threads, not millions.
  • A coroutine is scheduled in user space by your language runtime. Its stack starts tiny and grows on demand (kilobytes), and switching between coroutines is a handful of register saves that never leaves your process. Go’s goroutines, Python’s async tasks, and Kotlin’s coroutines are all this shape - cheap enough to spawn by the hundred thousand.

The cost gap is the whole reason coroutines exist. A wait that would pin an expensive thread instead parks a nearly-free coroutine.

Preemptive vs cooperative

The deeper split is who ends a task’s turn.

  • Threads are preempted. The OS interrupts a running thread on a timer and hands the core to another, whether the thread was ready to stop or not. No thread can hog the CPU forever just by never cooperating.
  • Coroutines yield. A coroutine keeps running until it reaches a point where it voluntarily gives up control - an await, a channel send, a blocking call the runtime knows how to intercept. Nothing forces it to stop.

That cooperative model is efficient but sharp-edged: a coroutine that enters a tight compute loop and never hits a yield point can starve every other coroutine sharing its thread. This is the classic event-loop stall - one slow function freezes the whole loop - and it is why a CPU-bound task in a cooperative runtime is a hazard (see CPU-bound vs IO-bound). Some runtimes file the edge down: modern Go can preempt a goroutine that runs too long, buying back some of the safety of real threads.

The two combine: M:N scheduling

Everything so far is language-agnostic: preemptive versus cooperative, and the cost of a single switch, are properties of any concurrency system, not of one runtime. In practice you rarely pick just one layer - a runtime stacks them. Many coroutines are multiplexed onto a small pool of OS threads, an M:N model, and a user-space scheduler keeps those threads busy: when a coroutine blocks on IO it is parked so another can run on the same thread, and a few threads keep thousands of coroutines moving.

That buys a real payoff and hides a real trap. The payoff: the number of OS threads is roughly your core count, so you write code as if concurrency were free while paying for only a handful of threads. The trap: those threads are shared, so a coroutine that blocks its underlying thread - on a syscall the runtime cannot intercept - can take the whole thread out of play and stall everything queued behind it. A serious M:N runtime therefore needs machinery to notice that and hand the waiting work off, and Go’s scheduler is the clearest worked example of getting it right. Its G-M-P model, work-stealing, and syscall handoff are enough of their own topic to live in a separate note: The Go scheduler: G, M, and P.

The same choice, four runtimes

Go picks M:N with a preemptive backstop. Every mainstream runtime is answering the same two questions - preemptive or cooperative, and how many OS threads - and they land in different places.

  • Java - real threads, and now coroutines too. Classic Java threads are 1:1: each maps to an OS thread the kernel preempts, so Java gets true multicore parallelism out of the box (this is the same shape as Go’s M). The cost is the usual one - a thread is heavy, so tens of thousands is a lot. Java 21’s virtual threads (Project Loom) add the coroutine layer on top: cheap user-space threads multiplexed onto a few carriers, the same M:N bet Go made, retrofitted onto a threaded runtime.
  • Python - threads that can’t run in parallel. CPython has the GIL, a single lock that lets only one thread execute Python bytecode at a time. The threading module gives you genuine OS threads, but the GIL serializes their CPU work, so they help with IO (a thread waiting on the network holds no lock) and do nothing for CPU-bound parallelism. The escapes are explicit: multiprocessing for real parallelism across separate processes, or asyncio for cooperative coroutines in one thread. (Python 3.13 ships an experimental free-threaded build that removes the GIL - the exception that proves how central it has been.)
  • JavaScript - one thread, and an event loop around it. Your JS runs on a single thread; there is no shared-memory threading in the language. What makes it feel concurrent is that IO is non-blocking - the runtime hands the wait to the OS and keeps pulling callbacks off the event loop, itself one more queue, so one thread juggles thousands of connections. Actual parallelism means a separate address space: Web Workers in the browser, worker_threads in Node, talking by messages, not shared memory.

The through-line: the kernel is the only thing that delivers real parallelism, and every runtime above either embraces its threads (Java, Go’s M), works around a lock that blocks them (Python), or leans entirely on non-blocking IO to avoid needing them (JavaScript).

Which one the work wants

The choice follows the workload, exactly as in CPU-bound vs IO-bound:

  • IO-bound, massively concurrent - tens of thousands of connections mostly waiting - wants coroutines. The waits are cheap, and you never spend a thread to sit idle.
  • CPU-bound, parallel - number crunching that must actually run at the same time - wants real threads on real cores. Coroutines multiplex work onto the cores you have; they do not conjure more of them.

The bottom line

Threads and coroutines are not rivals to pick between; they are layers that answer different questions. Threads are how work reaches a core and how the OS keeps one task from monopolizing it. Coroutines are how you keep tens of thousands of waits in flight without paying for a thread each. The question to keep straight is always the same: who owns the scheduler here, and what does a single wait cost?