GoByte Skills Episode 8 of 27, track Python (2 of 3)

The GIL: threads for waiting, processes for work

GoByte Skills #8: The GIL lets one thread run Python bytecode at a time. Four threads on CPU work finish no faster than one. Four processes really run in parallel. Measured, plus why the __main__ guard is not style.

Zoe is one of GoByte's characters. This post was drafted by AI agents in Zoe's voice, then fact checked, run and edited by the GoByte team.

Back to top

The GIL lets one thread run Python bytecode at a time, so threads help with I/O and processes help with CPU.

Animation for GoByte Skills #8: Threads take turns on the interpreter while processes run side by side.
Transcript

Four Python threads take turns holding a single interpreter lock while one bar fills at a time, then four processes each with their own lock fill their bars side by side.

import time
from concurrent.futures import (
    ThreadPoolExecutor, ProcessPoolExecutor)

def burn(n):
    while n:
        n -= 1

def run(Pool):
    t = time.perf_counter()
    with Pool(4) as ex:
        list(ex.map(burn, [10**8] * 4))
    return time.perf_counter() - t

if __name__ == "__main__":
    t = time.perf_counter()
    for _ in range(4):
        burn(10**8)
    print(f"serial    {time.perf_counter() - t:.2f}s")
    print(f"threads   {run(ThreadPoolExecutor):.2f}s")
    print(f"processes {run(ProcessPoolExecutor):.2f}s")

Measured#

One run on a 14 core Mac, CPython 3.14, load average about 5:

serial    3.11s
threads   3.17s
processes 0.85s

Threads stay within noise of serial. Four processes finish in roughly serial time divided by four, plus startup cost: 3.6x to 3.7x faster over three runs here. The numbers move with load; the shape does not.

Why#

CPython's global interpreter lock is one mutex per interpreter, and a thread must hold it to execute bytecode. Four CPU bound threads take turns on one interpreter. The holder is asked to let go every switch interval (sys.getswitchinterval(), 0.005 s by default), which buys fairness, not speed. Blocking calls such as time.sleep, socket reads and file I/O release the lock while they wait: map time.sleep over [1] * 4 with four threads and it takes about 1 s, not 4. Each process has its own interpreter and its own GIL, so processes run truly in parallel.

What processes cost#

Arguments and results are pickled through a pipe, and each worker starts a fresh interpreter. The default start method is spawn on macOS and, since 3.14, forkserver on Linux, so workers re-import your module. That is why the __main__ guard is mandatory, not style. Tiny tasks can be slower in a pool than inline.

Escape hatches#

C code can release the GIL around long loops, which is how some extensions use threads well. Since 3.13 there is an optional free threaded build (PEP 703); sys._is_gil_enabled() tells you which one you are on. It is not the default.

Rule of thumb#

Threads for waiting, processes for work. Four threads on CPU work is four cooks sharing one stove: very polite, same dinner time.

Report a mistake

Your product here? Partner with us

Back to top

Discussion

No comments yet. Signed in GoByte members with a verified e-mail can join. Community guidelines

Reading is open to everyone. Commenting and voting need a GoByte account with a verified e-mail.