GIL 不是 Python 的缺陷,而是 CPython 的一个实现取舍。搞清它保护什么、不保护什么,你才知道该用线程、进程还是协程。
你将学到
- GIL 到底是什么,以及为什么它只是 CPython 的实现细节,不是语言规范
- GIL 保护哪些东西、又完全不保护哪些东西
- 为什么 CPU 密集跑不快,而 IO 密集几乎不受影响
- 用一段代码亲眼复现
counter += 1的竞态 - 多线程 / 多进程 / asyncio / C 扩展四种模型的选型清单
前置知识
需要先理解线程、字节码的基本概念。建议先读 05 - 迭代器、生成器与协程,本篇是它的并发续集。
一、GIL 是什么:一句话和一段历史
GIL(Global Interpreter Lock,全局解释器锁)是一把互斥锁,它保证同一个 CPython 进程里,任意时刻只有一个线程在执行 Python 字节码。
划重点:GIL 是 CPython 的实现细节,不是 Python 语言特性。 语言规范里根本没有"全局锁"这个东西。只不过 CPython(你 python 命令默认用的那个解释器)因为历史原因选择了它。PyPy 有 GIL,Jython 和 IronPython 没有,而 Python 3.13 开始官方提供了可选的 free-threaded 版本。
import sys
import platform
print(sys.implementation.name) # 输出: cpython
print(platform.python_version()) # 输出: 3.11.x(或你本机的版本)
# 想验证"只有一个线程在执行字节码"?看下面这个例子
import threading
def spin():
while True: # 纯 Python 死循环 = 一直在占着 GIL
pass
# 开两个线程一起转,观察 CPU 占用
t1 = threading.Thread(target=spin, daemon=True)
t2 = threading.Thread(target=spin, daemon=True)
t1.start(); t2.start()
# 在任务管理器 / top 里你会看到:CPU 只跑到一个核(约 100%),
# 而不是两个核(约 200%)。这就是 GIL 的直观证据。
历史原因很简单:CPython 用引用计数(reference counting)管理内存。每个对象的 ob_refcnt 加一减一都不是原子操作,如果多个线程同时改,内存就会损坏。最省事的方案就是——拿一把大锁把所有字节码执行串起来。它让 CPython 的实现简单、C 扩展好写,代价就是多核 CPU 用不满。
二、GIL 保护什么,不保护什么
这是最容易被误解的地方。GIL 只保证单个字节码指令相对于其他线程是原子的,它不保证 Python 语句级别的原子性。
import threading
counter = 0
def worker():
global counter
for _ in range(100_000):
counter += 1 # 看起来是一行,其实是多步
# ---- 关键:counter += 1 到底做了什么? ----
# 它等价于下面三步(可以用 dis 模块验证):
import dis
dis.dis(worker)
# LOAD_GLOBAL counter <- 1. 读 counter
# LOAD_CONST 1
# BINARY_OP += 1 <- 2. 加 1
# STORE_GLOBAL counter <- 3. 写回 counter
# GIL 只保证上面每一条是原子的,不保证这三条打包在一起是原子的。
t1 = threading.Thread(target=worker)
t2 = threading.Thread(target=worker)
t1.start(); t2.start()
t1.join(); t2.join()
print(counter) # 输出: 大概率 < 200000,比如 187342
为什么丢了?线程 A 刚 LOAD_GLOBAL 读到 counter = 5,还没等写回,时间片到了,GIL 被切换给线程 B;B 也读到 5,加到 6 写回;A 回来把自己算的 6 写回。两次 +1 的结果只剩一次。这就是竞态条件(race condition)。
GIL 不保护的东西,远不止共享变量:
# GIL 不保护这些:
# 1. 多步操作(读-改-写)的复合原子性,如上
# 2. 数据结构在遍历时被其他线程修改
# 3. 文件、网络等外部资源的并发访问
# 4. 多个 C 扩展各自释放 GIL 后对同一块内存的访问
三、字节码切换:GIL 什么时候会被释放
CPython 不会让一个线程永远占着 GIL,它在两种时机切换:
- 时间片到期:从 Python 3.2 起用"抢占式 + 5 毫秒"的机制。执行了一定数量的字节码后,线程主动检查并可能让出 GIL。用
sys.setswitchinterval()可以调整这个间隔。 - 遇到会释放 GIL 的调用:这是最重要的一条。任何阻塞式 IO(读文件、
socket.recv、time.sleep、数据库查询)在进入系统调用前都会释放 GIL。 这就是 IO 密集为什么不受 GIL 困扰的原因。
import sys
# 把切换间隔调大,能更容易观察到"长时间不切换"
sys.setswitchinterval(0.1) # 每个线程最多连续持有 GIL 约 0.1 秒
print(sys.getswitchinterval()) # 输出: 0.1
# 反过来,sleep 是"友好"的:它主动放锁
import time
def io_like():
time.sleep(1) # 这 1 秒里,别的线程可以随便跑
# 所以 10 个 IO 线程各 sleep 1 秒,总耗时约 1 秒,不是 10 秒
结论一句话:GIL 惩罚的是"纯 Python 计算",放过的是"IO 等待"。
四、两个实验:CPU 密集 vs IO 密集
把同样的多线程代码跑在两件事上,结果天差地别。
import threading
import time
# ---------- 实验 1:CPU 密集 ----------
def cpu_bound(n):
total = 0
for i in range(n):
total += i * i
return total
N = 5_000_000
start = time.perf_counter()
cpu_bound(N)
cpu_bound(N)
serial_cpu = time.perf_counter() - start
start = time.perf_counter()
ts = [threading.Thread(target=cpu_bound, args=(N,)) for _ in range(2)]
for t in ts: t.start()
for t in ts: t.join()
thread_cpu = time.perf_counter() - start
print(f"CPU 串行: {serial_cpu:.3f}s")
print(f"CPU 双线程: {thread_cpu:.3f}s")
# 输出: CPU 串行: 1.20s
# 输出: CPU 双线程: 1.25s <- 几乎没变快!GIL 让两线程排队跑
# ---------- 实验 2:IO 密集 ----------
def io_bound():
time.sleep(0.5) # 模拟一次网络 / 磁盘等待
start = time.perf_counter()
io_bound(); io_bound()
serial_io = time.perf_counter() - start
start = time.perf_counter()
ts = [threading.Thread(target=io_bound) for _ in range(2)]
for t in ts: t.start()
for t in ts: t.join()
thread_io = time.perf_counter() - start
print(f"IO 串行: {serial_io:.3f}s")
print(f"IO 双线程: {thread_io:.3f}s")
# 输出: IO 串行: 1.001s
# 输出: IO 双线程: 0.501s <- 直接减半!等待期间 GIL 是放开的
这两个实验,是你以后所有并发选型决策的证据基础。别背结论,记住现象。
五、free-threaded:3.13 的 no-GIL 版本
Python 3.13 开始,官方提供了一个实验性的 free-threaded(无 GIL)构建。它不是默认版本,而是一个单独的 python3.13t 可执行文件(PEP 703)。
import sys
# 判断当前解释器是不是 free-threaded 构建
print(sys._is_gil_enabled()) if hasattr(sys, "_is_gil_enabled") else None
# free-threaded 版本输出: True(表示 GIL 已禁用)
# 普通版本可能没有这个属性
# 在 free-threaded 构建里,多线程真的能并行跑纯 Python 代码了
但要清醒:
- 它是实验性的,单线程性能目前比带 GIL 版本慢一截(写时复制 / 偏向引用计数等弥补手段有开销)。
- 大量 C 扩展还没适配,遇到没适配的扩展 GIL 会被自动重新打开。
- 现阶段它解决的是"终于有路",而不是"生产可用"。
所以 2026 年写生产代码,仍然按"有 GIL"来设计。free-threaded 是未来,不是现在的答案。
六、选型:四张牌怎么打
到了真正要写并发代码的时候,先问自己一个问题:瓶颈是等,还是算?
| 场景 | 首选 | 原因 |
|---|---|---|
| CPU 密集(数值计算、图像处理、加解密) | 多进程 multiprocessing |
每个进程独立 GIL,真能占满多核 |
| IO 密集(爬虫、API 调用、文件读写) | 多线程 或 asyncio | GIL 在等待时释放,并发即有效 |
| 海量 IO 连接(几千个 socket) | asyncio | 单线程事件循环,无上下文切换开销 |
| 性能敏感的计算热点 | C 扩展 / Cython / Rust | 在 C 层释放 GIL,真正的并行 |
再补充几条经验法则:
# 场景 A:CPU 密集 -> 多进程
from concurrent.futures import ProcessPoolExecutor
with ProcessPoolExecutor() as pool:
results = list(pool.map(cpu_bound_task, data))
# 4 核机器上,4 个进程 ≈ 4 倍加速(通信开销除外)
# 场景 B:IO 密集且并发不高 -> 线程池
from concurrent.futures import ThreadPoolExecutor
with ThreadPoolExecutor(max_workers=20) as pool:
results = list(pool.map(fetch_url, urls))
# 4096 * 20 个线程也没问题,线程很轻,等待时都在放锁
# 场景 C:IO 密集且连接巨大 -> asyncio
import asyncio, aiohttp
async def fetch_all(urls):
async with aiohttp.ClientSession() as s:
tasks = [s.get(u) for u in urls]
return await asyncio.gather(*tasks)
# 单线程就能扛住上万连接,见 08 篇
# 场景 D:热点在纯计算 -> 换算法 / 上 C
# 例如用 numpy 向量化替代 Python 循环,
# numpy 的重运算在 C 层会主动释放 GIL,多线程也能加速
注意一个反直觉的坑:多进程不是免费的。 进程启动要几百毫秒,进程间传输数据要序列化(pickle)。如果每个任务只跑几十毫秒,用多进程反而更慢。任务越"重"、数据越"小",多进程越划算。
常见坑
import threading
# ❌ 坑 1:以为"加锁"是多余的操作,共享计数不加锁
counter = 0
def bad():
global counter
for _ in range(100_000):
counter += 1 # 竞态:结果会丢数
# ✅ 正解:用锁把复合操作包起来
lock = threading.Lock()
def good():
global counter
for _ in range(100_000):
with lock:
counter += 1 # 读-改-写整体原子
# 另一个更省事的正解:直接用 threading 不擅长的场景换队列,
# 或用 multiprocessing 的共享内存计数器(见 07 篇)
# ❌ 坑 2:CPU 密集用多线程,然后抱怨"怎么没变快"
import threading
def crunch():
sum(i * i for i in range(10_000_000))
ts = [threading.Thread(target=crunch) for _ in range(8)]
for t in ts: t.start()
for t in ts: t.join()
# 8 个线程抢一把 GIL,可能比自己顺序跑还慢(多了切换成本)
# ✅ 正解:CPU 密集换进程池
from concurrent.futures import ProcessPoolExecutor
with ProcessPoolExecutor(max_workers=8) as pool:
list(pool.map(lambda _: crunch(), range(8)))
# 真正并行,8 核就是约 8 倍(前提是你的任务确实重)
# ❌ 坑 3:把 GIL 当成"线程安全"的护身符
import threading
shared = {}
def add(k):
shared[k] = shared.get(k, 0) + 1 # 又是读-改-写,GIL 救不了你
# ✅ 正解:要么加锁,要么用 collections.defaultdict 配合单键原子操作,
# 要么干脆让每个线程算自己的局部结果,最后再合并(无共享即无竞态)
小结
- GIL 是 CPython 的实现细节,保障引用计数等内部状态一致,不是 Python 语言规范。
- 它只保证单条字节码原子,不保证
counter += 1这类多步操作原子——竞态照旧存在。 - GIL 在时间片到期和进入阻塞 IO 时释放,所以 IO 密集不怕 GIL,CPU 密集很怕。
- 选型口诀:CPU 密集用多进程,IO 密集用线程或 asyncio,热点计算下推到 C/库。
- 多进程有启动和序列化开销,任务太轻时反而更慢;别把"能并行"当成"一定更快"。
- free-threaded(3.13+)是未来方向,当前仍按带 GIL 设计生产代码。
延伸阅读
- PEP 703 —— Making the Global Interpreter Lock Optional
- CPython 源码
ceval.c里的CALL/PyEval_EvalFrameDefault主循环 - 《CPython Internals》关于 GIL 与引用计数的章节
文章回复
0 条公开回复