The Go scheduler: G, M, and P

Budding

5 min read

Goroutines feel free. You spawn them by the hundred thousand and the language behaves as if that costs nothing. It does not cost nothing - it costs a scheduler, and Go’s is worth understanding because it is a clean, visible answer to a hard problem: how do you run an enormous number of cheap concurrent tasks on the small number of OS threads a machine can actually afford?

This is Go’s take on M:N scheduling - many coroutines multiplexed onto few threads. The general model, and why a runtime wants it, is in Threads vs coroutines: who decides to yield; this note is how Go implements it.

Three roles: G, M, P

Go’s runtime splits the work across three things:

  • G - the goroutine. The unit of work: a function plus its own stack, about 2 KB to start and grown on demand. Creating one is cheap, so having hundreds of thousands alive at once is ordinary.
  • M - the machine. A real OS thread, the only thing that can actually execute on a core. Threads are heavy (on the order of a megabyte), so the runtime keeps a pool and reuses them instead of asking the OS for a fresh one per goroutine.
  • P - the processor. A logical execution slot that carries the permission to run Go code plus a local run queue of ready goroutines. There are GOMAXPROCS of them, defaulting to the number of cores, so the count of Ps is the real ceiling on how much Go code runs in parallel.

A goroutine runs only when an M holds a P and picks a G off that P’s local queue. That piece in the middle - the P - is the whole trick: a fixed number of run permits that a thread must acquire before it can execute Go code. It is what caps parallelism and what lets each core keep its own queue without a global lock.

flowchart TB
  subgraph P1["P (local run queue)"]
    G1(G) --- G2(G) --- G3(G)
  end
  subgraph P2["P (local run queue)"]
    G4(G) --- G5(G)
  end
  P1 --> M1[M · OS thread]
  P2 --> M2[M · OS thread]
  M1 --> C1[CPU core]
  M2 --> C2[CPU core]
  GRQ["global run queue<br/>(overflow + newly woken Gs)"] -.-> P1
  GRQ -.-> P2

Each P drains its own local queue; a global run queue holds the overflow and newly woken goroutines that no P has claimed yet.

M and P are not the same thing

The easiest trap in this model is the word “processor.” A P is not a CPU core. It is a logical run-permit Go invented, and there are exactly GOMAXPROCS of them regardless of what the hardware looks like. The cores belong to the OS; Go never schedules onto them directly, it schedules onto Ps and lets the OS place the threads on whatever cores exist.

That separation is the whole point, because the two counts move independently:

  • P is a fixed budget - exactly GOMAXPROCS slots, each with a queue, capping how many goroutines run Go code at the same instant.
  • M is a flexible pool - the runtime spins up threads as it needs them, and there are usually more Ms than Ps. A goroutine blocked in a syscall keeps its own M but releases its P (the handoff below), so a program with GOMAXPROCS=4 and a thousand goroutines stuck on disk reads can have close to a thousand Ms alive while only four of them run Go code.

What keeps it efficient

Three mechanics do the real work:

  • Work-stealing. A P that drains its local queue does not idle. It steals half the goroutines from a P that still has a backlog. Each P balances itself, with no central coordinator and no global lock, so no core sits hungry while another is buried.
  • The syscall handoff. When a goroutine makes a blocking syscall the runtime cannot turn into a park-and-resume, its M is stuck in the kernel until the call returns. Rather than freeze everything queued behind it, Go detaches the P from that stuck M and pairs it with a spare M, so the rest of that queue keeps running. One blocking call costs a thread, not a whole core’s worth of goroutines. This single mechanic is what separates a robust M:N runtime from a naive one.
  • Preemption. Goroutines are cooperative by default - they yield at channel operations, blocking calls, and function-call safepoints - but a background monitor (sysmon) watches for one that has run too long, around 10 ms, and forces it to yield. That preemptive backstop means a single tight compute loop can slow a P but cannot starve it forever (the hazard in CPU-bound vs IO-bound).

Why it matters

You can treat goroutines as nearly free and let the scheduler place them - but the abstraction leaks in exactly the spots the model predicts. A CPU-bound goroutine competes for one of only GOMAXPROCS slots. A C library that blocks its thread ties up an M. And GOMAXPROCS is a knob worth checking when Go runs inside a container, where the core count it sees may not match what the platform actually grants. Knowing the G, the M, and the P is what turns those from surprises into behavior you can predict.