The Go scheduler: G, M, and P
Goroutines feel free. You spawn them by the hundred thousand and the language behaves as if that costs nothing. It does not cost nothing - it costs a scheduler, and Go’s is worth understanding because it is a clean, visible answer to a hard problem: how do you run an enormous number of cheap concurrent tasks on the small number of OS threads a machine can actually afford?
This is Go’s take on M:N scheduling - many coroutines multiplexed onto few threads. The general model, and why a runtime wants it, is in Threads vs coroutines: who decides to yield; this note is how Go implements it.
Three roles: G, M, P
Go’s runtime splits the work across three things:
- G - the goroutine. The unit of work: a function plus its own stack, about 2 KB to start and grown on demand. Creating one is cheap, so having hundreds of thousands alive at once is ordinary.
- M - the machine. A real OS thread, the only thing that can actually execute on a core. Threads are heavy (on the order of a megabyte), so the runtime keeps a pool and reuses them instead of asking the OS for a fresh one per goroutine.
- P - the processor. A logical execution slot that carries the permission to run Go code plus a local run queue of ready goroutines. There are
GOMAXPROCSof them, defaulting to the number of cores, so the count ofPs is the real ceiling on how much Go code runs in parallel.
A goroutine runs only when an M holds a P and picks a G off that P’s local queue. That piece in the middle - the P - is the whole trick: a fixed number of run permits that a thread must acquire before it can execute Go code. It is what caps parallelism and what lets each core keep its own queue without a global lock.
flowchart TB
subgraph P1["P (local run queue)"]
G1(G) --- G2(G) --- G3(G)
end
subgraph P2["P (local run queue)"]
G4(G) --- G5(G)
end
P1 --> M1[M · OS thread]
P2 --> M2[M · OS thread]
M1 --> C1[CPU core]
M2 --> C2[CPU core]
GRQ["global run queue<br/>(overflow + newly woken Gs)"] -.-> P1
GRQ -.-> P2
Each P drains its own local queue; a global run queue holds the overflow and newly woken goroutines that no P has claimed yet.
M and P are not the same thing
The easiest trap in this model is the word “processor.” A P is not a CPU core. It is a logical run-permit Go invented, and there are exactly GOMAXPROCS of them regardless of what the hardware looks like. The cores belong to the OS; Go never schedules onto them directly, it schedules onto Ps and lets the OS place the threads on whatever cores exist.
That separation is the whole point, because the two counts move independently:
- P is a fixed budget - exactly
GOMAXPROCSslots, each with a queue, capping how many goroutines run Go code at the same instant. - M is a flexible pool - the runtime spins up threads as it needs them, and there are usually more Ms than Ps. A goroutine blocked in a syscall keeps its own M but releases its
P(the handoff below), so a program withGOMAXPROCS=4and a thousand goroutines stuck on disk reads can have close to a thousand Ms alive while only four of them run Go code.
What keeps it efficient
Three mechanics do the real work:
- Work-stealing. A
Pthat drains its local queue does not idle. It steals half the goroutines from aPthat still has a backlog. EachPbalances itself, with no central coordinator and no global lock, so no core sits hungry while another is buried. - The syscall handoff. When a goroutine makes a blocking syscall the runtime cannot turn into a park-and-resume, its
Mis stuck in the kernel until the call returns. Rather than freeze everything queued behind it, Go detaches thePfrom that stuckMand pairs it with a spareM, so the rest of that queue keeps running. One blocking call costs a thread, not a whole core’s worth of goroutines. This single mechanic is what separates a robust M:N runtime from a naive one. - Preemption. Goroutines are cooperative by default - they yield at channel operations, blocking calls, and function-call safepoints - but a background monitor (
sysmon) watches for one that has run too long, around 10 ms, and forces it to yield. That preemptive backstop means a single tight compute loop can slow aPbut cannot starve it forever (the hazard in CPU-bound vs IO-bound).
Why it matters
You can treat goroutines as nearly free and let the scheduler place them - but the abstraction leaks in exactly the spots the model predicts. A CPU-bound goroutine competes for one of only GOMAXPROCS slots. A C library that blocks its thread ties up an M. And GOMAXPROCS is a knob worth checking when Go runs inside a container, where the core count it sees may not match what the platform actually grants. Knowing the G, the M, and the P is what turns those from surprises into behavior you can predict.
Related
- Threads vs coroutines: who decides to yield - the general threads-vs-coroutines model this implements
- CPU-bound vs IO-bound - which workloads the scheduler carries and which it cannot
- Garbage collection: the pause you didn't schedule - the same runtime that schedules your goroutines also collects their garbage
- How Programs Actually Run MOC - the runtime layer this note belongs to
- Performance & Latency MOC - the map these notes hang from