Python 学习笔记
#语言 · #Python · #ThinkPython · #PythonCookbook
基于 Allen Downey Think Python(第 3 版)和 Python Cookbook(第 3 版)系统学习 Python。从"程序如何运行"的基本心智模型开始,到高级数据结构和工程实践,建立对 Python 的完整认知。
一、程序之道(Think Python Ch1-4)
1.1 Python 的心智模型
Python 程序执行 = 解释器逐条执行语句
关键概念:
- 变量 = 对对象的引用(标签贴在对象上),而非存储值的容器
- 对象有类型(type),变量没有类型
- 赋值 a = b:让 a 也引用 b 指向的对象(没有复制)python
# Python 的引用语义
a = [1, 2, 3] # a → 列表对象 [1,2,3]
b = a # b → 同一个列表对象(没有复制!)
b.append(4)
print(a) # [1, 2, 3, 4] ← a 也变了
# 可变 vs 不可变
x = "hello" # str 不可变
y = x # y → 同一个 str 对象
y = y + " world" # y → 新 str 对象,x 不变
print(x) # "hello"Python 引用语义图解:变量是对象的"标签",而非"容器"。
mermaid
graph LR
subgraph Vars["栈上变量"]
A["a"]
B["b"]
X["x"]
Y["y"]
end
subgraph Heap["堆中对象"]
L1["list [1,2,3,4]<br/>id=0x7f01"]
S1["str 'hello'<br/>id=0x7f02"]
S2["str 'hello world'<br/>id=0x7f03"]
end
A --> L1
B --> L1
X --> S1
Y --> S2
S1 -.-|"不可变,修改=新建"| S2
style L1 fill:#FF6B6B,color:#fff
style S1 fill:#4ECDC4,color:#fff
style S2 fill:#4ECDC4,color:#fff关键理解:
a = [1,2,3]不是"把列表放入变量 a",而是"创建标签 a,贴在列表对象上"。这导致同对象修改会"穿透"到所有引用者。C 语言中变量就是存储位置本身,Go 中除指针外也是值拷贝——Python 的模型与它们根本不同。
可变共享的经典翻车场景:
python
matrix = [[0] * 3] * 3 # 看似 3×3,实际三个子列表是同一个对象
matrix[0][0] = 1
print(matrix) # [[1,0,0],[1,0,0],[1,0,0]] ← 三行全变!
# 正确写法: [[0]*3 for _ in range(3)]1.2 类型与值
| 类型 | 可变性 | 示例 | 说明 |
|---|---|---|---|
int | 不可变 | 42 | 任意精度(无溢出),对象是小整数缓存池(-5~256) |
float | 不可变 | 3.14 | IEEE 754 双精度 64bit,0.1+0.2 != 0.3 |
str | 不可变 | "hello" | Unicode 字符串,底层 PyUnicodeObject,紧凑存储(ASCII→1B) |
bytes | 不可变 | b"\x00\xff" | 原始字节序列 |
bool | 不可变 | True/False | 是 int 的子类:True==1, False==0 |
None | — | None | 单例,is None 而非 == None |
list | 可变 | [1, 2, 3] | 动态数组(array of PyObject*),扩容 1.125x |
tuple | 不可变 | (1, 2) | 不可变列表,但元素本身可以是可变对象 |
dict | 可变 | {"a":1} | 开放寻址哈希表(3.6+ 保证插入有序) |
set | 可变 | {1, 2} | 开放寻址哈希表,与 dict 共享底层实现 |
1.3 控制流
python
# if/elif/else 链
if x < 0:
sign = "negative"
elif x == 0:
sign = "zero"
else:
sign = "positive"
# for 遍历可迭代对象(不直接操作索引)
for item in items:
process(item)
for i, item in enumerate(items):
print(f"{i}: {item}")
# while(少用,for 更 Pythonic)
while condition:
update()
# match/case(3.10+,模式匹配)
match status:
case 200:
print("OK")
case 404:
print("Not Found")
case _:
print("Unknown")1.4 函数设计(Think Python Ch3-4)
python
# 参数类型
def func(required, optional="default", *args, **kwargs):
"""参数:位置实参 → 关键字实参 → *args → **kwargs"""
pass
# 参数传递:全是传引用(pass-by-reference)
def append_item(lst):
lst.append(4) # 修改了原对象
def reassign(lst):
lst = [1, 2, 3] # 只是改变了局部变量引用,不影响调用者
# 函数是"一等公民"
def add(a, b): return a + b
multiply = lambda a, b: a * b # lambda:单表达式匿名函数
operations = [add, multiply]
result = operations[0](3, 5) # 8二、数据结构(Think Python Ch9-13 + Cookbook Ch1)
2.1 字符串与文本处理
Python 3 的字符串是 Unicode 原生支持的——这也是它和 Python 2 最根本的区别之一。str 类型在内部使用 PyUnicodeObject 存储,根据字符串中最大字符自动选择最紧凑的编码(1字节/ASCII、2字节/UCS2 或 4字节/UCS4)。这意味着纯英文文本每个字符占 1 字节,中文每个字符占 2~4 字节。
与 Go 的重要差异:Go 的 string 底层是 []byte 只读视图(总是 UTF-8),用 len() 得到字节数而非字符数;Python 的 len(s) 始终返回字符数(Unicode 码点数),这让文本处理更直观。
不可变性:Python 字符串和 Go 一样不可变。每次
s + "!"都创建新对象。高频拼接务必用''.join()或io.StringIO。
Python 的字符串 API 设计遵循"直觉优先"原则——方法与自然语言描述一致(如 startswith 而非 starts_with),返回值也一目了然:
python
# str 核心操作
s = "hello world"
s.startswith("he") # True
s.endswith("ld") # True
s.find("wo") # 6(未找到返回 -1)
s.replace("world", "Python") # "hello Python"
s.split() # ["hello", "world"]
",".join(["a", "b"]) # "a,b"
s.strip() # 去两端空白
# 格式化
f"{name}: {value:.2f}" # f-string(3.6+,最快)
"{}: {:.2f}".format(name, value)
"%s: %.2f" % (name, value)三种格式化方式各有适用场景:f-string 适合简单变量嵌入(最快),.format() 适合模板复用,% 格式化是 Python 2 遗留但仍在日志场景广泛使用(logging 模块的惰性求值依赖它)。
2.2 列表与序列
Python 的 list 是最常用的数据结构,底层是 PyListObject——一个动态扩容的 PyObject* 指针数组。这意味着 list 可以容纳任意类型的元素混合存放(因为存的是指针而非值本身),但也意味着每个元素的访问需要额外的指针跳转。
扩容策略:newsize ≈ oldsize + (oldsize >> 3) + 3(约 1.125x)。这比 C++ std::vector(2x)保守得多,也比 Go slice(<256 时 2x)慢热——Python 追求节约内存而非激进扩容。
列表推导式是 Python 的标志性语法,它将"映射+过滤"逻辑压缩到一行,背后是 C 级别的循环优化(比普通 for + append 快约 30-50%)。但注意:推导式的结果是立刻计算的完整列表——大数据量时应使用生成器表达式。
切片是 Python 最具表达力的序列操作。它的语法 [start:stop:step] 看似简单,但能优雅地解决"取子序列"、"反转"、"跳跃取样"等几乎所有常见需求。理解切片的左闭右开规则(start 包含,stop 不包含)是关键——这与 range() 的设计一脉相承:
python
# 切片:list[start:stop:step](都是整数,左闭右开)
nums = [0, 1, 2, 3, 4, 5]
nums[1:4] # [1, 2, 3]
nums[::-1] # [5, 4, 3, 2, 1, 0] 反序
# 列表推导式(Pythonic 核心)
squares = [x**2 for x in range(10) if x % 2 == 0]
# 等价于:
squares = []
for x in range(10):
if x % 2 == 0:
squares.append(x**2)
# 嵌套推导
matrix = [[i * j for j in range(3)] for i in range(3)]
# 生成器表达式(惰性求值,省内存)
sum(x**2 for x in range(10**6))推导式 vs map/filter:Python 社区推崇推导式而非
map()/filter(),因为前者更可读。[x for x in data if x > 0]比list(filter(lambda x: x > 0, data))直观得多。Guido 甚至曾提议从 Python 3 中移除map/filter,虽未实施,但态度明确。
2.3 字典(dict)
Python 3.6+ 的 dict 有两个关键设计:开放寻址(解决哈希冲突)和紧凑存储(entries 与 indices 分离,保证插入顺序)。这使 dict 在内存效率和迭代性能之间取得了极好的平衡。
紧凑字典内部结构:
mermaid
flowchart LR
subgraph indices["indices (稀疏哈希表)"]
I0["[0] 空"]
I1["[1] →0"]
I2["[2] 空"]
I3["[3] →1"]
I4["[4] 空"]
I5["[5] 空"]
I6["[6] →2"]
I7["[7] 空"]
end
subgraph entries["entries (紧凑数组, 插入顺序)"]
direction TB
E0["[0] hash('a'), 'a', 1"]
E1["[1] hash('b'), 'b', 2"]
E2["[2] hash('c'), 'c', 3"]
end
I1 --> E0
I3 --> E1
I6 --> E2结构说明:
indices是稀疏哈希表,存储entries数组的下标。entries是紧凑数组,按插入顺序存放键值对。查找 key 时先哈希 → indices 下标 → entries 位置 → 比较 key 值。删除不立即回收空间,而是标记为"dummy"——这就是为什么频繁增删 dict 后建议调用gc.collect()。
mermaid
flowchart TB
subgraph Find["查找 key='b'"]
A["hash('b') → 0x..." ] --> B["hash & mask → indices下标"]
B --> C{"indices[下标] 有值?"}
C -->|"有"| D["取 entries[indices[i]]"]
D --> E{"key == 'b'?"}
E -->|"是"| F["返回 value=2"]
E -->|"否"| G["探测下一个位置"]
G --> C
C -->|"无"| H["KeyError"]
end
style F fill:#4CAF50,color:#fff
style H fill:#FF6B6B,color:#fff核心权衡:为什么使用开放寻址而非拉链法(如 Java HashMap)?
| 方面 | 开放寻址(Python) | 拉链法(Java/Go) |
|---|---|---|
| 缓存局部性 | ✅ 好(数组连续) | ❌ 差(链表跳转) |
| 内存碎片 | ✅ 低 | ❌ 节点分散 |
| 删除代价 | ⚠️ 需标记位(dummy) | ✅ 简单 |
| 负载因子 | ~0.67 | ~0.75 |
| 删除操作频繁 | ⚠️ dummy 积累需重建 | ✅ |
python
# CPython dict 底层实现(3.6+):
# 开放寻址哈希表 + 紧凑键值对数组
# 结构:entries = [[hash, key_ptr, value_ptr], ...]
# indices = [entry_index or DKIX_EMPTY, ...]
# 插入顺序 = 键值对在 entries 中的顺序 → 天然有序!
# 负载因子:2/3 时触发扩容
# 与 Go map 对比:
# Cookbook 1.6-1.7:defaultdict 和 OrderedDict
from collections import defaultdict, OrderedDict, Counter
d = defaultdict(list) # 键不存在时自动创建 list
d["a"].append(1)
# 计数
words = ["a", "b", "a", "c", "b", "a"]
counts = Counter(words) # Counter({"a": 3, "b": 2, "c": 1})
counts.most_common(2) # [("a", 3), ("b", 2)]
# 字典运算(Cookbook 1.8-1.9)
prices = {"AAPL": 150.0, "GOOG": 2800.0, "MSFT": 330.0}
min_price = min(zip(prices.values(), prices.keys())) # (150.0, "AAPL")
sorted_prices = sorted(zip(prices.values(), prices.keys()))2.4 集合(set)
python
# Python set 与 dict 共享底层(开放寻址哈希表),只存 key 不存 value
# 对比 C++ set:红黑树(有序);Java HashSet:HashMap 包装
# Python set:无序(3.6+ 实际插入顺序),O(1) 查找
a = {1, 2, 3, 4}
b = {3, 4, 5, 6}
a | b # 并集 {1,2,3,4,5,6}
a & b # 交集 {3,4}
a - b # 差集 {1,2}
a ^ b # 对称差 {1,2,5,6}2.5 元组与序列解包
python
# Cookbook 1.1-1.3:解包
p = ("Alice", 30, "Engineer")
name, age, job = p
# 星号解包
first, *middle, last = [1, 2, 3, 4, 5] # first=1, middle=[2,3,4], last=5
# namedtuple(不可变轻量类)
from collections import namedtuple
Point = namedtuple("Point", ["x", "y"])
p = Point(3, 4)
p.x, p.y # 3, 4
p._asdict() # {"x": 3, "y": 4}
# 对比 dataclass(3.7+):可变,支持默认值和方法
from dataclasses import dataclass
@dataclass
class Point:
x: float
y: float
def distance(self) -> float:
return (self.x**2 + self.y**2) ** 0.5三、面向对象设计(Think Python Ch15-18)
Python 的面向对象与 Java/C++ 有本质区别。Java 强调"一切皆对象,必须定义类",Python 更倾向于"需要时才用类"。一个简单的脚本可以用函数搞定,中途发现需要状态时再重构为类——这种渐进式设计是 Python OOP 的核心哲学。
最关键的差异:Python 不需要声明接口(interface),因为它是鸭子类型。一个对象只要"走起来像鸭子、叫起来像鸭子",它就是鸭子——不需要显式 implements Duck。这使得 Python 的代码更灵活,但也要求开发者更加自律(文档和测试来弥补类型系统的宽松)。
以下是 Python 面向对象的核心概念,从基本类定义到元类。
3.1 类与对象
python
class Card:
"""一张扑克牌"""
suit_names = ["Clubs", "Diamonds", "Hearts", "Spades"]
rank_names = [None, "Ace", "2", "3", ..., "King"]
def __init__(self, suit: int, rank: int):
self.suit = suit # 实例属性
self.rank = rank
def __str__(self) -> str:
return f"{self.rank_names[self.rank]} of {self.suit_names[self.suit]}"
def __lt__(self, other: "Card") -> bool:
return self.rank < other.rank # 支持排序3.2 继承与多态
python
# 鸭子类型:"如果它走路像鸭子,叫起来像鸭子,它就是鸭子"
# 不需要显式继承某个基类,只需要有对应的方法
class Duck:
def quack(self): return "Quack!"
class Person:
def quack(self): return "I'm imitating a duck!"
def make_quack(thing):
print(thing.quack()) # 任何有 quack() 方法的对象都能传入
make_quack(Duck()) # Quack!
make_quack(Person()) # I'm imitating a duck!
# 抽象基类(ABC):当需要严格约束时
from abc import ABC, abstractmethod
class Quacker(ABC):
@abstractmethod
def quack(self) -> str: ...3.3 魔术方法
| 方法 | 触发 | 用途 |
|---|---|---|
__init__ | obj = Class() | 初始化 |
__str__ | str(obj), print(obj) | 人类可读 |
__repr__ | repr(obj), 交互环境 | 调试输出,尽量 eval(repr(x)) == x |
__eq__ | a == b | 值相等比较 |
__lt__ | a < b | 排序比较 |
__len__ | len(obj) | 长度 |
__getitem__ | obj[key] | 索引/切片 |
__iter__ | for x in obj | 迭代 |
__call__ | obj() | 可调用对象 |
__enter__/__exit__ | with obj | 上下文管理器 |
3.4 属性与描述符
python
# @property:把方法伪装成属性
class Circle:
def __init__(self, radius: float):
self._radius = radius
@property
def radius(self) -> float:
return self._radius
@radius.setter
def radius(self, value: float):
if value < 0:
raise ValueError("radius must be >= 0")
self._radius = value
@property
def area(self) -> float:
return 3.14159 * self._radius ** 2
c = Circle(5)
print(c.area) # 78.53975(像属性一样访问)
c.radius = 10 # 触发 setter四、迭代器与生成器(Think Python Ch19)
生成器是 Python 最强大的特性之一,它的核心思想是惰性求值——不是一次性计算所有结果存入内存,而是"按需生产"。这对处理大规模数据流至关重要。
理解生成器的关键:yield 关键字将普通函数变为"可暂停的函数"。每次调用 next() 时,生成器从上次 yield 处恢复执行,产生下一个值后再次暂停。这种"协程"思想后来发展为 async/await(第九章)。
与 Go 的对比:Go 用 channel + goroutine 实现生产者-消费者模式,Python 用生成器。两者都是"流式处理",但 Go 是并发的多 goroutine 通信,Python 生成器是单线程协作式——更简单但无法利用多核。
4.1 迭代器协议
python
# for x in obj 等价于:
it = iter(obj) # 调用 obj.__iter__() 返回迭代器
while True:
try:
x = next(it) # 调用 it.__next__()
except StopIteration:
break4.2 生成器(yield)
python
# 生成器函数:包含 yield → 返回生成器对象
def fibonacci(n: int):
"""生成前 n 个斐波那契数(惰性求值)"""
a, b = 0, 1
for _ in range(n):
yield a
a, b = b, a + b
# 生成器是单向的:只能产出值,不能接收值(除非用 send())
# 生成器表达式
squares = (x**2 for x in range(10**6)) # 不立刻分配内存
sum(x**2 for x in range(10**6)) # 直接传给函数五、迭代器与生成器技巧(Cookbook Ch4)
5.1 手动消费迭代器
python
items = [1, 2, 3]
it = iter(items)
next(it, None) # 1(默认值 None 防止 StopIteration)
next(it, None) # 25.2 itertools 模块
python
from itertools import islice, chain, zip_longest, product, combinations
# 对迭代器切片
first_5 = list(islice(iterator, 5))
# 链式迭代
for x in chain([1, 2], "ab", {3, 4}):
print(x) # 1 2 a b 3 4
# 排列组合
list(product("AB", "12")) # [('A','1'),('A','2'),('B','1'),('B','2')]
list(combinations("ABC", 2)) # [('A','B'),('A','C'),('B','C')]六、文件 I/O 与数据序列化(Cookbook Ch5-6)
Python 的文件 I/O 设计体现了"简洁优于复杂"的哲学。with 语句自动管理资源生命周期(即使发生异常也会关闭文件),比 Java 的 try-with-resources 更早出现,比 Go 的 defer f.Close() 更安全(defer 只在函数返回时执行,with 在退出代码块时立即执行)。
什么是"Pythonic"的文件操作:用 for line in file 替代 while True: line = f.readline(),前者惰性逐行读取,不会把整个文件加载到内存。对大文件(GB 级别),这是唯一正确的方式。
6.1 文件读写的 Pythonic 方式
python
# with 自动关闭文件
with open("data.txt", "r", encoding="utf-8") as f:
for line in f: # 惰性逐行读取,不占内存
process(line)
# 一次读全部
with open("data.txt") as f:
text = f.read() # 适合小文件
# 写文件
with open("output.txt", "w") as f:
f.write("hello\n")
print("world", file=f) # print 写入文件6.2 JSON / CSV / Pickle
python
import json, csv
# JSON
data = json.loads('{"name": "Alice", "age": 30}')
json.dumps(data, indent=2, ensure_ascii=False)
# CSV
with open("data.csv") as f:
reader = csv.DictReader(f)
for row in reader:
print(row["column_name"])
# Pickle(Python 专有,不能跨语言)
import pickle
pickle.dumps(obj) # 序列化为 bytes
pickle.loads(data) # 反序列化七、函数进阶(Cookbook Ch7)
7.1 可调用对象
python
# 自定义可调用对象(比闭包更清晰)
class Countdown:
def __init__(self, start: int):
self.start = start
def __call__(self) -> bool:
self.start -= 1
return self.start > 0
cd = Countdown(5)
while cd():
print("working...") # 执行 4 次7.2 偏函数与函数组合
python
from functools import partial, reduce
# 偏函数:固定部分参数
parse_int_base2 = partial(int, base=2)
parse_int_base2("1010") # 10
# reduce:累积计算
total = reduce(lambda a, b: a * b, [1, 2, 3, 4]) # 247.3 lambda vs 具名函数
python
# lambda:仅用于简单表达式
sorted(users, key=lambda u: u.age)
# 如果逻辑变复杂 → 提取为函数
def get_fullname(user):
return f"{user.last_name} {user.first_name}"
sorted(users, key=get_fullname)八、类与对象进阶(Cookbook Ch8-9)
8.1 字符串表示
python
class Point:
def __repr__(self): return f"Point({self.x!r}, {self.y!r})"
# !r 表示用 repr() 格式化,区别:
# f"{s}" → 调用 str(s)
# f"{s!r}" → 调用 repr(s)8.2 创建管理型上下文管理器
python
from contextlib import contextmanager
@contextmanager
def tempdir():
"""创建临时目录,自动清理"""
import tempfile, shutil
path = tempfile.mkdtemp()
try:
yield path
finally:
shutil.rmtree(path)
with tempdir() as d:
open(f"{d}/file.txt", "w").write("hello")
# 自动清理8.3 __slots__ 节约内存
python
class Point:
__slots__ = ["x", "y"] # 禁止创建 __dict__,固定属性
def __init__(self, x, y):
self.x = x
self.y = y
# 节省内存场景:创建数百万个简单对象时
# 常规类:每个实例有 __dict__(至少 64 字节 + 每个属性)
# __slots__:只有固定槽位(每个槽位 8 字节指针)8.4 元编程:装饰器
python
from functools import wraps
def log_call(func):
"""装饰器:记录每次调用"""
@wraps(func) # 保留原函数元信息(__name__, __doc__ 等)
def wrapper(*args, **kwargs):
print(f"calling {func.__name__}({args}, {kwargs})")
return func(*args, **kwargs)
return wrapper
@log_call
def add(a, b):
"""Return the sum of a and b."""
return a + b
add(3, 5) # calling add((3, 5), {}) → 8
help(add) # "Return the sum of a and b."(wraps 保留了 docstring)九、并发编程(Cookbook Ch12 + 深入)
9.1 GIL 原理
GIL 是 CPython 的"阿喀琉斯之踵"——它让多线程在多核 CPU 上对 CPU 密集型任务完全失效。理解它不是"坏设计",而是引用计数式内存管理的历史代价。
mermaid
sequenceDiagram
participant T1 as 线程1 (CPU密集)
participant GIL as GIL
participant T2 as 线程2 (CPU密集)
participant OS as 操作系统
OS->>T1: 调度到 CPU 核0
T1->>GIL: 获取 GIL ✅
T1->>T1: 执行 100 条字节码
T1->>GIL: 释放 GIL
Note over GIL,OS: 线程1被OS剥夺CPU
OS->>T2: 调度到 CPU 核1
T2->>GIL: 获取 GIL ✅
T2->>T2: 执行 100 条字节码
T2->>GIL: 释放 GIL
Note over GIL,OS: ← 两个线程轮流执行,但绝不同时!
Note over T1,T2: 多核CPU上,实际只有一个核在工作<br/>线程切换反而增加开销GIL 关键事实:
- CPython 解释器级别的互斥锁,每个 Python 进程只有一个
- 同一时刻只有一个线程能执行 Python 字节码
- 原因:CPython 的引用计数不是线程安全的,GIL 是保护它的"大锁"
- 后果:
· CPU 密集型多线程 → 无效(甚至比单线程慢,因为有切换开销)
· I/O 密集型多线程 → 有效(I/O 等待时主动释放 GIL)
- 绕过方式:多进程(multiprocessing)、C 扩展释放 GIL、用 PyPy/Jython(无 GIL)
- Python 3.13 引入"Free-threaded CPython"(实验性,可禁用 GIL)text
GIL 的底层实现:
CPython 源码 (Python/ceval_gil.h):
- 本质是一个非递归互斥锁 (PyMutex)
- 每执行 100 条字节码 (sys.getswitchinterval() ≈ 5ms) 或遇到 I/O 操作时释放
- 释放后立即重新尝试获取(其他线程有机会抢到)
- 在 POSIX 上使用 pthread_mutex + pthread_cond 实现条件等待
字节码执行循环 (ceval.c):
while (1) {
take_gil(); // 获取 GIL
while (ticks-- > 0) {
opcode = *next_instr++;
switch (opcode) {
case LOAD_FAST: ...
case BINARY_ADD: ...
// 每个 opcode 执行后检查是否需要释放 GIL
if (ticks == 0) break; // 时间片用完 → 释放 GIL
}
}
drop_gil(); // 释放 GIL → 其他线程有机会
}
引用计数为什么需要 GIL:
PyObject->ob_refcnt 的 ++/-- 不是原子的(在大多数平台)。
如果两个线程同时修改同一个对象的 ob_refcnt → 计数错乱
→ 内存泄漏或提前释放(UAF)。GIL 保证了每次只有一个人在改。
释放 GIL 的时机(由字节码解释器控制):
- 执行了 100 条字节码(默认,可用 sys.setswitchinterval() 调整)
- 调用 C 扩展中显式释放 GIL 的函数(Py_BEGIN_ALLOW_THREADS)
- I/O 操作(read/write/connect/accept 等,CPython 在进入阻塞前释放 GIL)
- time.sleep() — 主动让出
为什么 I/O 多线程有效? → I/O 阻塞时 GIL 释放,其他线程获取 GIL 继续执行
为什么 CPU 多线程无效? → 没有 I/O 释放点,只能等 100 条字节码配额用完asyncio 事件循环的底层分析
text
asyncio 本质上是在单线程中模拟并发——它不是真正的并行。
事件循环的核心循环 (简化版):
while True:
1. 检查就绪的 Future/Task(epoll_wait / select)
2. 从就绪队列取一个 Task → 执行 Task.__step()
3. Task 内部遇到 await → 挂起当前 Task → 注册回调
4. 切换到下一个就绪 Task → 重复
await 到底做了什么?
await asyncio.sleep(1) 等价于:
1. 创建一个 TimerHandle(1秒后触发)
2. 当前协程 yield 控制权 → 返回事件循环
3. 1 秒后 TimerHandle 触发 → 事件循环恢复该协程
与多线程的本质区别:
- 协程切换: 用户态,无系统调用,~50ns
- 线程切换: 内核态,保存/恢复上下文,~1-10μs
- 但协程是协作式的: 一个协程死循环 → 整个事件循环卡死!python
# asyncio 的微观示例 — 理解 await 挂起和恢复
import asyncio
async def probe():
print(f"[0ms] probe start")
await asyncio.sleep(0.01) # ← 这里发生了什么?
# 1. sleep(0.01) 返回一个 Future
# 2. await 将当前协程"挂起"到该 Future
# 3. 事件循环切换到其他就绪协程
# 4. 0.01 秒后 Future 就绪 → 事件循环恢复此协程
print(f"[10ms] probe resumed")mermaid
sequenceDiagram
participant Loop as 事件循环
participant C1 as 协程1
participant C2 as 协程2
participant Timer as 定时器
Loop->>C1: 恢复协程1
C1->>C1: 执行字节码...
C1->>Timer: await asyncio.sleep(0.01)
Note over C1: 挂起, yield 给事件循环
Loop->>C2: 恢复协程2
C2->>C2: 执行字节码...
Timer-->>Loop: 0.01s 后触发
Loop->>C1: 协程1 恢复执行并发模型决策树:
mermaid
flowchart TD
Q["你的任务是?"]
Q -->|"I/O 密集型<br/>(网络/文件/DB)"| IO["多线程<br/>ThreadPoolExecutor"]
Q -->|"CPU 密集型<br/>(计算/图像)"| CPU["多进程<br/>ProcessPoolExecutor"]
Q -->|"高并发 I/O<br/>(10000+连接)"| ASYNC["asyncio<br/>异步协程"]
Q -->|"I/O + CPU 混合"| MIX["多进程 + 每进程内 asyncio"]
IO --> IO2["✅ GIL 在 IO 时释放<br/>适合 Web 请求/爬虫"]
CPU --> CPU2["✅ 绕过 GIL<br/>每个进程独立 GIL"]
ASYNC --> ASYNC2["✅ 单线程协作式<br/>极高的 I/O 吞吐"]
MIX --> MIX2["✅ 最佳并行<br/>各进程绑定 CPU+协程"]
style IO2 fill:#4CAF50,color:#fff
style CPU2 fill:#2196F3,color:#fff
style ASYNC2 fill:#FF9800,color:#fff
style MIX2 fill:#9C27B0,color:#fff9.2 多线程(I/O 密集型)
python
from concurrent.futures import ThreadPoolExecutor, as_completed
def fetch_url(url: str) -> str:
import urllib.request
return urllib.request.urlopen(url).read().decode()
urls = ["https://httpbin.org/get"] * 10
with ThreadPoolExecutor(max_workers=5) as executor:
futures = {executor.submit(fetch_url, url): url for url in urls}
for future in as_completed(futures):
result = future.result()9.3 多进程(CPU 密集型)
python
from concurrent.futures import ProcessPoolExecutor
def compute_heavy(n: int) -> int:
return sum(i * i for i in range(n))
with ProcessPoolExecutor(max_workers=4) as executor:
results = list(executor.map(compute_heavy, [10**7] * 8))9.4 asyncio(高并发 I/O)
asyncio 的核心是事件循环(Event Loop):一个单线程循环,管理着成百上千个协程,在它们之间来回切换。当协程遇到 await 时主动让出控制权,事件循环立即切换到另一个就绪的协程——这就是"协作式多任务"。
mermaid
sequenceDiagram
participant EL as 事件循环 (单线程)
participant C1 as 协程1 (fetch url1)
participant C2 as 协程2 (fetch url2)
participant IO as 网络 I/O
EL->>C1: 执行到 await session.get()
C1-->>EL: 让出控制权(yield)
EL->>IO: 注册 fd1 到 epoll
EL->>C2: 调度协程2
C2->>C2: 执行到 await session.get()
C2-->>EL: 让出控制权
EL->>IO: 注册 fd2 到 epoll
Note over EL: 无协程可调度,等待 I/O
IO-->>EL: fd1 就绪!(数据到达)
EL->>C1: 恢复协程1
C1->>C1: 处理响应数据
C1-->>EL: 完成
IO-->>EL: fd2 就绪!
EL->>C2: 恢复协程2
C2->>C2: 处理响应数据
C2-->>EL: 完成
Note over EL: 全部协程完成,退出关键对比:asyncio vs Go goroutine
| asyncio (Python) | goroutine (Go) | |
|---|---|---|
| 调度方式 | 协作式(await 显式让出) | 抢占式(runtime 自动切换) |
| 阻塞后果 | 忘记 await = 阻塞整个事件循环 | goroutine 阻塞只挂起自身 |
| CPU 利用 | 单核 | 多核 |
| 内存开销 | 极小(协程栈 KB 级) | 小(goroutine 栈 2KB 起) |
| 学习成本 | 中等(理解 await/event loop) | 低(go 关键字即可) |
核心原则:asyncio 中绝对不能调用阻塞函数(如
time.sleep()),必须使用await asyncio.sleep()。否则整个事件循环会被卡住,所有协程全部停滞——这是新手最容易犯的错误。
python
import asyncio
async def fetch(session, url: str) -> str:
async with session.get(url) as resp:
return await resp.text()
async def main():
import aiohttp
async with aiohttp.ClientSession() as session:
tasks = [fetch(session, f"https://httpbin.org/get?id={i}")
for i in range(100)]
results = await asyncio.gather(*tasks)
asyncio.run(main())9.5 GIL vs Go
| 维度 | Python (CPython) | Go |
|---|---|---|
| 并发模型 | 线程 + GIL / asyncio 协程 | goroutine(抢占式) |
| 并行 | 多进程(有开销) | 多核并行(原生) |
| 阻塞语义 | await 显式让出 | goroutine 自动切换 |
| 学习曲线 | 需要理解 await/event loop | goroutine/channel 简洁 |
十、元编程(Cookbook Ch9)
10.1 类装饰器
python
def add_repr(cls):
"""自动为类添加 __repr__ 方法"""
def __repr__(self):
attrs = ", ".join(f"{k}={v!r}" for k, v in self.__dict__.items()
if not k.startswith("_"))
return f"{cls.__name__}({attrs})"
cls.__repr__ = __repr__
return cls
@add_repr
class Person:
def __init__(self, name: str, age: int):
self.name = name
self.age = age
p = Person("Alice", 30)
print(p) # Person(name='Alice', age=30)10.2 元类
python
# 元类 = 类的类(type 是默认元类)
# 场景:ORM(Django/SQLAlchemy)、接口强制执行
class ValidateFields(type):
def __new__(cls, name, bases, attrs):
for key, value in attrs.items():
if key.startswith("__"):
continue
if not hasattr(value, "__annotations__"):
raise TypeError(f"{key} must have type annotation")
return super().__new__(cls, name, bases, attrs)
class Base(metaclass=ValidateFields):
pass十一、测试与调试(Think Python App B + Cookbook Ch14)
11.1 doctest
python
def add(a: int, b: int) -> int:
"""
>>> add(1, 2)
3
>>> add(-1, 1)
0
"""
return a + b
# python -m doctest module.py11.2 unittest / pytest
python
import pytest
def divide(a, b):
if b == 0:
raise ValueError("division by zero")
return a / b
def test_divide():
assert divide(10, 2) == 5.0
with pytest.raises(ValueError):
divide(10, 0)
@pytest.mark.parametrize("a,b,expected", [(10, 2, 5), (9, 3, 3)])
def test_parametrized(a, b, expected):
assert divide(a, b) == expected11.3 logging
python
import logging
logging.basicConfig(level=logging.INFO,
format="%(asctime)s [%(levelname)s] %(name)s: %(message)s")
logger = logging.getLogger(__name__)
logger.info("Starting process")
logger.warning("Low disk space")
logger.error("Connection failed", exc_info=True) # 自动包含异常堆栈十二、标准库精选(Cookbook 各章)
| 模块 | 用途 | 示例 |
|---|---|---|
pathlib | 面向对象路径操作 | Path("a/b").read_text() |
subprocess | 子进程管理 | subprocess.run(["ls", "-l"], capture_output=True) |
argparse | 命令行参数解析 | parser.add_argument("--verbose", action="store_true") |
tempfile | 临时文件/目录 | tempfile.NamedTemporaryFile() |
shutil | 高级文件操作 | shutil.copy2(src, dst) |
datetime | 日期时间 | datetime.now(tz=timezone.utc) |
collections | 高级容器 | deque, Counter, ChainMap |
functools | 高阶函数 | lru_cache, partial, reduce |
hashlib | 哈希摘要 | hashlib.sha256(data).hexdigest() |
re | 正则表达式 | re.search(r"\d+", text) |
argparse 完整示例
python
import argparse
parser = argparse.ArgumentParser(description="Process some files")
parser.add_argument("files", nargs="+", help="input files")
parser.add_argument("-o", "--output", default="out.txt", help="output file")
parser.add_argument("-v", "--verbose", action="store_true", help="verbose output")
args = parser.parse_args(["-v", "-o", "result.txt", "a.txt", "b.txt"])
# args.files = ["a.txt", "b.txt"], args.output = "result.txt", args.verbose = True十三、第三方库
13.1 FastAPI(Web 框架)
python
from fastapi import FastAPI
from pydantic import BaseModel
app = FastAPI()
class Item(BaseModel):
name: str
price: float
@app.get("/items/{item_id}")
async def get_item(item_id: int) -> Item:
return Item(name="Widget", price=19.99)
@app.post("/items")
async def create_item(item: Item) -> Item:
# 自动校验、序列化
return item13.2 常用库速查
| 库 | 用途 | import |
|---|---|---|
requests | HTTP 客户端 | requests.get(url) |
aiohttp | 异步 HTTP | async with session.get(url) |
pydantic | 数据校验 | BaseModel |
pandas | 数据分析 | pd.read_csv("data.csv") |
numpy | 数值计算 | np.array([1, 2, 3]) |
SQLAlchemy | ORM | 声明式映射 |
Celery | 分布式任务队列 | @app.task |
十四、Python 内存管理深度
Python 的内存管理是"引用计数 + 分代 GC"的双层机制——引用计数处理大部分对象的释放,分代 GC 兜底处理循环引用:
mermaid
flowchart TB
subgraph RefCount["引用计数 (实时)"]
RC1["每次赋值/传参: ob_refcnt++"]
RC2["每次离开作用域: ob_refcnt--"]
RC3["refcnt==0 → 立即释放 (tp_dealloc)"]
RC4["✅ 延迟低, 内存释放及时"]
RC5["❌ 循环引用无法释放!"]
end
subgraph GenGC["分代 GC (周期性)"]
GC1["第0代: 新对象,频繁扫描"]
GC2["第1代: 存活过的对象"]
GC3["第2代: 长期存活对象,很少扫描"]
GC4["用三色标记找到不可达的循环引用"]
end
RC5 -->|"触发 gc.collect()"| GC1循环引用实战:
python
# 循环引用:内存泄漏的经典场景
class Node:
def __init__(self, name):
self.name = name
self.ref = None
a = Node("A")
b = Node("B")
a.ref = b # a → b
b.ref = a # b → a ← 循环引用!
# a 和 b 的 refcnt 永远 ≥ 1 → 引用计数不会释放
# 需要分代 GC 扫描才能找到并回收CPython vs PyPy vs Cython — 性能对比
| 维度 | CPython (3.12) | PyPy (3.10) | Cython |
|---|---|---|---|
| JIT 编译 | ❌ | ✅ 追踪 JIT | ❌ (AOT 编译为 C 扩展) |
| 纯 Python 性能 | 1x (基准) | 3-5x | N/A |
| C 扩展兼容 | ✅ 全兼容 | 🟡 部分兼容(C API 慢) | ✅ 生成 C 扩展 |
| GIL | ✅ 有 (3.13 可选无 GIL) | ❌ 无 GIL | 取决于 C 代码 |
| 内存占用 | 中等 | 高(JIT 编译缓存) | 低(接近 C) |
| 启动速度 | ✅ 极快 | ❌ 慢(预热 JIT) | ✅ 快 |
| 适用 | 通用 Python 代码 | 纯 Python 计算密集型 | 需要 C 速度的 Python |
PyPy 不是万能加速器:它加速的是纯 Python 循环和计算(如 for 循环中的数值计算)。如果你的代码大量使用 C 扩展库(numpy/pandas)或 I/O 密集型——PyPy 不会带来明显提升。Cython 适合将某个 CPU 热点函数编译为 C 扩展。
十五、Python 与其他语言对比
| 特性 | Python | Go | C++ | Java |
|---|---|---|---|---|
| 类型系统 | 动态(鸭子类型) | 静态 | 静态 | 静态 |
| 并发模型 | GIL + asyncio | goroutine | 线程 + 锁 | 线程 + Executor |
| 内存管理 | 引用计数 + 分代 GC | GC(并发标记-清除) | RAII / 手动 | GC |
| 编译 | 解释 → 字节码 | 编译 → 原生 | 编译 → 原生 | 编译 → 字节码 |
| dict 实现 | 开放寻址(有序) | 链地址法(无序) | std::unordered_map | HashMap |
| list 扩容 | ~1.125x | <256→2x, ≥256→1.25x | 2x (GCC) | 1.5x |
| 适合场景 | 脚本/AI/Web后端 | 微服务/基础设施 | 系统/游戏/高频 | 企业级后端 |
登录后即可发表评论 👇