System calls: the border crossing
A running program is walled off from the machine it runs on. It cannot read a file, send a packet, or start another process by itself, because those touch hardware and other programs, and the operating system reserves them. The only way through the wall is a system call: a controlled request that crosses from your code into the kernel, has the kernel do the privileged thing on your behalf, and returns. Every read, write, open, socket, and mmap is one of these crossings, and the crossing itself is not free.
Two modes, one door
The CPU runs at one of two privilege levels. User mode, where your code lives, cannot execute privileged instructions or touch arbitrary memory and devices. Kernel mode can. A system call is a deliberate, narrow doorway between them: a special instruction that traps into the kernel at a fixed entry point, switches to kernel mode, runs vetted kernel code, and switches back with a result. The wall is the whole point - it is what stops one process from reading another’s memory or halting the machine. The syscall is the single audited gate in it.
Why the crossing costs
A syscall looks like a function call in your source, but underneath it is a different animal. Crossing the boundary means:
- switching privilege mode and entering the kernel at its trap handler,
- saving and restoring register state around the transition,
- often polluting the CPU caches and TLB with kernel work, so your code resumes a little colder,
- and, if the call blocks - a disk read, a socket with no data yet - a full context switch to some other thread while you wait.
None of that happens on an ordinary function call. The rough orders of magnitude: a plain call is a nanosecond or less, a syscall is hundreds of nanoseconds to a few microseconds, and a real disk or network operation is orders of magnitude beyond that. The crossing is cheap next to real IO and ruinously expensive next to a function call, which is why doing millions of tiny syscalls is a classic performance sink.
The craft is fewer, fatter crossings
Almost every fast-IO technique is the same move: amortize the crossing over more work.
- Buffering. Standard library IO collects many small writes and flushes them in one syscall. Writing a line at a time only costs one crossing per line if you forgot the buffer.
- Vectored IO.
readvandwritevmove several buffers in a single call instead of one apiece. - Readiness multiplexing.
epollandkqueuelet one syscall watch thousands of sockets, so a server does not cross the boundary per connection to ask “anything yet?” - Submission queues.
io_uringlets a program queue many operations and reap their results with very few crossings, pushing the boundary cost toward zero per operation. - Mapping instead of calling.
mmapmaps a file into your address space so you read it as memory, skipping the per-chunkreadcalls entirely.
The bridge to concurrency
The syscall boundary is where the concurrency story and the runtime story meet. A blocking syscall is precisely what pins a thread for the length of a wait, which is the cost async was built to dodge: async runtimes route IO through non-blocking syscalls plus a readiness interface like epoll, so one thread can supervise thousands of in-flight operations (see Async/await: syntax over a state machine). It is also why Go’s scheduler needs its syscall handoff - a goroutine stuck in a blocking call would otherwise take its whole thread down with it (see The Go scheduler: G, M, and P). Whether a wait costs you a thread comes down to which kind of syscall you made.
The bottom line
A system call is the price of leaving your own process, and you cannot opt out of it - the wall is what keeps the system safe. What you control is how often you pay. Slow IO code is rarely slow because the kernel is slow; it is slow because it crosses the boundary too many times, a function call’s worth of logic wrapped around a microsecond’s worth of toll, millions of times over.
Related
- The Go scheduler: G, M, and P - the syscall handoff that keeps one blocking call from freezing a core
- Async/await: syntax over a state machine - non-blocking syscalls are what make cooperative IO possible
- CPU-bound vs IO-bound - IO-bound work is work dominated by these crossings and the waits behind them
- Queues Are Everywhere - epoll and io_uring are the kernel handing you queues
- How Programs Actually Run MOC - the map these runtime notes hang from