Skip to content

Python 学习笔记 ​

#语言 · #Python · #ThinkPython · #PythonCookbook

基于 Allen Downey Think Python(第 3 版)和 Python Cookbook(第 3 版)系统学习 Python。从"程序如何运行"的基本心智模型开始,到高级数据结构和工程实践,建立对 Python 的完整认知。


一、程序之道(Think Python Ch1-4) ​

1.1 Python 的心智模型 ​

Python 程序执行 = 解释器逐条执行语句

关键概念:
- 变量 = 对对象的引用(标签贴在对象上),而非存储值的容器
- 对象有类型(type),变量没有类型
- 赋值 a = b:让 a 也引用 b 指向的对象(没有复制)
python
# Python 的引用语义
a = [1, 2, 3]   # a → 列表对象 [1,2,3]
b = a            # b → 同一个列表对象(没有复制!)
b.append(4)
print(a)         # [1, 2, 3, 4] ← a 也变了

# 可变 vs 不可变
x = "hello"      # str 不可变
y = x            # y → 同一个 str 对象
y = y + " world" # y → 新 str 对象,x 不变
print(x)         # "hello"

Python 引用语义图解:变量是对象的"标签",而非"容器"。

mermaid
graph LR
    subgraph Vars["栈上变量"]
        A["a"]
        B["b"]
        X["x"]
        Y["y"]
    end
    subgraph Heap["堆中对象"]
        L1["list [1,2,3,4]<br/>id=0x7f01"]
        S1["str 'hello'<br/>id=0x7f02"]
        S2["str 'hello world'<br/>id=0x7f03"]
    end
    A --> L1
    B --> L1
    X --> S1
    Y --> S2
    S1 -.-|"不可变,修改=新建"| S2

    style L1 fill:#FF6B6B,color:#fff
    style S1 fill:#4ECDC4,color:#fff
    style S2 fill:#4ECDC4,color:#fff

关键理解:a = [1,2,3] 不是"把列表放入变量 a",而是"创建标签 a,贴在列表对象上"。这导致同对象修改会"穿透"到所有引用者。C 语言中变量就是存储位置本身,Go 中除指针外也是值拷贝——Python 的模型与它们根本不同。

可变共享的经典翻车场景:

python
matrix = [[0] * 3] * 3   # 看似 3×3,实际三个子列表是同一个对象
matrix[0][0] = 1
print(matrix)  # [[1,0,0],[1,0,0],[1,0,0]] ← 三行全变!
# 正确写法: [[0]*3 for _ in range(3)]

1.2 类型与值 ​

类型可变性示例说明
int不可变42任意精度(无溢出),对象是小整数缓存池(-5~256)
float不可变3.14IEEE 754 双精度 64bit,0.1+0.2 != 0.3
str不可变"hello"Unicode 字符串,底层 PyUnicodeObject,紧凑存储(ASCII→1B)
bytes不可变b"\x00\xff"原始字节序列
bool不可变True/False是 int 的子类:True==1, False==0
None—None单例,is None 而非 == None
list可变[1, 2, 3]动态数组(array of PyObject*),扩容 1.125x
tuple不可变(1, 2)不可变列表,但元素本身可以是可变对象
dict可变{"a":1}开放寻址哈希表(3.6+ 保证插入有序)
set可变{1, 2}开放寻址哈希表,与 dict 共享底层实现

1.3 控制流 ​

python
# if/elif/else 链
if x < 0:
    sign = "negative"
elif x == 0:
    sign = "zero"
else:
    sign = "positive"

# for 遍历可迭代对象(不直接操作索引)
for item in items:
    process(item)

for i, item in enumerate(items):
    print(f"{i}: {item}")

# while(少用,for 更 Pythonic)
while condition:
    update()

# match/case(3.10+,模式匹配)
match status:
    case 200:
        print("OK")
    case 404:
        print("Not Found")
    case _:
        print("Unknown")

1.4 函数设计(Think Python Ch3-4) ​

python
# 参数类型
def func(required, optional="default", *args, **kwargs):
    """参数:位置实参 → 关键字实参 → *args → **kwargs"""
    pass

# 参数传递:全是传引用(pass-by-reference)
def append_item(lst):
    lst.append(4)       # 修改了原对象

def reassign(lst):
    lst = [1, 2, 3]     # 只是改变了局部变量引用,不影响调用者

# 函数是"一等公民"
def add(a, b): return a + b
multiply = lambda a, b: a * b   # lambda:单表达式匿名函数
operations = [add, multiply]
result = operations[0](3, 5)    # 8

二、数据结构(Think Python Ch9-13 + Cookbook Ch1) ​

2.1 字符串与文本处理 ​

Python 3 的字符串是 Unicode 原生支持的——这也是它和 Python 2 最根本的区别之一。str 类型在内部使用 PyUnicodeObject 存储,根据字符串中最大字符自动选择最紧凑的编码(1字节/ASCII、2字节/UCS2 或 4字节/UCS4)。这意味着纯英文文本每个字符占 1 字节,中文每个字符占 2~4 字节。

与 Go 的重要差异:Go 的 string 底层是 []byte 只读视图(总是 UTF-8),用 len() 得到字节数而非字符数;Python 的 len(s) 始终返回字符数(Unicode 码点数),这让文本处理更直观。

不可变性:Python 字符串和 Go 一样不可变。每次 s + "!" 都创建新对象。高频拼接务必用 ''.join() 或 io.StringIO。

Python 的字符串 API 设计遵循"直觉优先"原则——方法与自然语言描述一致(如 startswith 而非 starts_with),返回值也一目了然:

python
# str 核心操作
s = "hello world"
s.startswith("he")          # True
s.endswith("ld")            # True
s.find("wo")               # 6(未找到返回 -1)
s.replace("world", "Python")  # "hello Python"
s.split()                   # ["hello", "world"]
",".join(["a", "b"])       # "a,b"
s.strip()                   # 去两端空白

# 格式化
f"{name}: {value:.2f}"     # f-string(3.6+,最快)
"{}: {:.2f}".format(name, value)
"%s: %.2f" % (name, value)

三种格式化方式各有适用场景:f-string 适合简单变量嵌入(最快),.format() 适合模板复用,% 格式化是 Python 2 遗留但仍在日志场景广泛使用(logging 模块的惰性求值依赖它)。

2.2 列表与序列 ​

Python 的 list 是最常用的数据结构,底层是 PyListObject——一个动态扩容的 PyObject* 指针数组。这意味着 list 可以容纳任意类型的元素混合存放(因为存的是指针而非值本身),但也意味着每个元素的访问需要额外的指针跳转。

扩容策略:newsize ≈ oldsize + (oldsize >> 3) + 3(约 1.125x)。这比 C++ std::vector(2x)保守得多,也比 Go slice(<256 时 2x)慢热——Python 追求节约内存而非激进扩容。

列表推导式是 Python 的标志性语法,它将"映射+过滤"逻辑压缩到一行,背后是 C 级别的循环优化(比普通 for + append 快约 30-50%)。但注意:推导式的结果是立刻计算的完整列表——大数据量时应使用生成器表达式。

切片是 Python 最具表达力的序列操作。它的语法 [start:stop:step] 看似简单,但能优雅地解决"取子序列"、"反转"、"跳跃取样"等几乎所有常见需求。理解切片的左闭右开规则(start 包含,stop 不包含)是关键——这与 range() 的设计一脉相承:

python
# 切片:list[start:stop:step](都是整数,左闭右开)
nums = [0, 1, 2, 3, 4, 5]
nums[1:4]    # [1, 2, 3]
nums[::-1]   # [5, 4, 3, 2, 1, 0]  反序

# 列表推导式(Pythonic 核心)
squares = [x**2 for x in range(10) if x % 2 == 0]
# 等价于:
squares = []
for x in range(10):
    if x % 2 == 0:
        squares.append(x**2)

# 嵌套推导
matrix = [[i * j for j in range(3)] for i in range(3)]

# 生成器表达式(惰性求值,省内存)
sum(x**2 for x in range(10**6))

推导式 vs map/filter:Python 社区推崇推导式而非 map()/filter(),因为前者更可读。[x for x in data if x > 0] 比 list(filter(lambda x: x > 0, data)) 直观得多。Guido 甚至曾提议从 Python 3 中移除 map/filter,虽未实施,但态度明确。

2.3 字典(dict) ​

Python 3.6+ 的 dict 有两个关键设计:开放寻址(解决哈希冲突)和紧凑存储(entries 与 indices 分离,保证插入顺序)。这使 dict 在内存效率和迭代性能之间取得了极好的平衡。

紧凑字典内部结构:

mermaid
flowchart LR
    subgraph indices["indices (稀疏哈希表)"]
        I0["[0] 空"]
        I1["[1] →0"]
        I2["[2] 空"]
        I3["[3] →1"]
        I4["[4] 空"]
        I5["[5] 空"]
        I6["[6] →2"]
        I7["[7] 空"]
    end

    subgraph entries["entries (紧凑数组, 插入顺序)"]
        direction TB
        E0["[0] hash('a'), 'a', 1"]
        E1["[1] hash('b'), 'b', 2"]
        E2["[2] hash('c'), 'c', 3"]
    end

    I1 --> E0
    I3 --> E1
    I6 --> E2

结构说明:indices 是稀疏哈希表,存储 entries 数组的下标。entries 是紧凑数组,按插入顺序存放键值对。查找 key 时先哈希 → indices 下标 → entries 位置 → 比较 key 值。删除不立即回收空间,而是标记为"dummy"——这就是为什么频繁增删 dict 后建议调用 gc.collect()。

mermaid
flowchart TB
    subgraph Find["查找 key='b'"]
        A["hash('b') → 0x..." ] --> B["hash & mask → indices下标"]
        B --> C{"indices[下标] 有值?"}
        C -->|"有"| D["取 entries[indices[i]]"]
        D --> E{"key == 'b'?"}
        E -->|"是"| F["返回 value=2"]
        E -->|"否"| G["探测下一个位置"]
        G --> C
        C -->|"无"| H["KeyError"]
    end
    style F fill:#4CAF50,color:#fff
    style H fill:#FF6B6B,color:#fff

核心权衡:为什么使用开放寻址而非拉链法(如 Java HashMap)?

方面开放寻址(Python)拉链法(Java/Go)
缓存局部性✅ 好(数组连续)❌ 差(链表跳转)
内存碎片✅ 低❌ 节点分散
删除代价⚠️ 需标记位(dummy)✅ 简单
负载因子~0.67~0.75
删除操作频繁⚠️ dummy 积累需重建✅
python
# CPython dict 底层实现(3.6+):
#   开放寻址哈希表 + 紧凑键值对数组
#   结构:entries = [[hash, key_ptr, value_ptr], ...]
#         indices = [entry_index or DKIX_EMPTY, ...]
#   插入顺序 = 键值对在 entries 中的顺序 → 天然有序!
#   负载因子:2/3 时触发扩容

# 与 Go map 对比:

# Cookbook 1.6-1.7:defaultdict 和 OrderedDict
from collections import defaultdict, OrderedDict, Counter

d = defaultdict(list)       # 键不存在时自动创建 list
d["a"].append(1)

# 计数
words = ["a", "b", "a", "c", "b", "a"]
counts = Counter(words)     # Counter({"a": 3, "b": 2, "c": 1})
counts.most_common(2)       # [("a", 3), ("b", 2)]

# 字典运算(Cookbook 1.8-1.9)
prices = {"AAPL": 150.0, "GOOG": 2800.0, "MSFT": 330.0}
min_price = min(zip(prices.values(), prices.keys()))  # (150.0, "AAPL")
sorted_prices = sorted(zip(prices.values(), prices.keys()))

2.4 集合(set) ​

python
# Python set 与 dict 共享底层(开放寻址哈希表),只存 key 不存 value
# 对比 C++ set:红黑树(有序);Java HashSet:HashMap 包装
# Python set:无序(3.6+ 实际插入顺序),O(1) 查找

a = {1, 2, 3, 4}
b = {3, 4, 5, 6}
a | b   # 并集 {1,2,3,4,5,6}
a & b   # 交集 {3,4}
a - b   # 差集 {1,2}
a ^ b   # 对称差 {1,2,5,6}

2.5 元组与序列解包 ​

python
# Cookbook 1.1-1.3:解包
p = ("Alice", 30, "Engineer")
name, age, job = p

# 星号解包
first, *middle, last = [1, 2, 3, 4, 5]  # first=1, middle=[2,3,4], last=5

# namedtuple(不可变轻量类)
from collections import namedtuple
Point = namedtuple("Point", ["x", "y"])
p = Point(3, 4)
p.x, p.y  # 3, 4
p._asdict()  # {"x": 3, "y": 4}

# 对比 dataclass(3.7+):可变,支持默认值和方法
from dataclasses import dataclass
@dataclass
class Point:
    x: float
    y: float
    def distance(self) -> float:
        return (self.x**2 + self.y**2) ** 0.5

三、面向对象设计(Think Python Ch15-18) ​

Python 的面向对象与 Java/C++ 有本质区别。Java 强调"一切皆对象,必须定义类",Python 更倾向于"需要时才用类"。一个简单的脚本可以用函数搞定,中途发现需要状态时再重构为类——这种渐进式设计是 Python OOP 的核心哲学。

最关键的差异:Python 不需要声明接口(interface),因为它是鸭子类型。一个对象只要"走起来像鸭子、叫起来像鸭子",它就是鸭子——不需要显式 implements Duck。这使得 Python 的代码更灵活,但也要求开发者更加自律(文档和测试来弥补类型系统的宽松)。

以下是 Python 面向对象的核心概念,从基本类定义到元类。

3.1 类与对象 ​

python
class Card:
    """一张扑克牌"""
    suit_names = ["Clubs", "Diamonds", "Hearts", "Spades"]
    rank_names = [None, "Ace", "2", "3", ..., "King"]

    def __init__(self, suit: int, rank: int):
        self.suit = suit   # 实例属性
        self.rank = rank

    def __str__(self) -> str:
        return f"{self.rank_names[self.rank]} of {self.suit_names[self.suit]}"

    def __lt__(self, other: "Card") -> bool:
        return self.rank < other.rank   # 支持排序

3.2 继承与多态 ​

python
# 鸭子类型:"如果它走路像鸭子,叫起来像鸭子,它就是鸭子"
# 不需要显式继承某个基类,只需要有对应的方法

class Duck:
    def quack(self): return "Quack!"

class Person:
    def quack(self): return "I'm imitating a duck!"

def make_quack(thing):
    print(thing.quack())  # 任何有 quack() 方法的对象都能传入

make_quack(Duck())    # Quack!
make_quack(Person())  # I'm imitating a duck!

# 抽象基类(ABC):当需要严格约束时
from abc import ABC, abstractmethod
class Quacker(ABC):
    @abstractmethod
    def quack(self) -> str: ...

3.3 魔术方法 ​

方法触发用途
__init__obj = Class()初始化
__str__str(obj), print(obj)人类可读
__repr__repr(obj), 交互环境调试输出,尽量 eval(repr(x)) == x
__eq__a == b值相等比较
__lt__a < b排序比较
__len__len(obj)长度
__getitem__obj[key]索引/切片
__iter__for x in obj迭代
__call__obj()可调用对象
__enter__/__exit__with obj上下文管理器

3.4 属性与描述符 ​

python
# @property:把方法伪装成属性
class Circle:
    def __init__(self, radius: float):
        self._radius = radius

    @property
    def radius(self) -> float:
        return self._radius

    @radius.setter
    def radius(self, value: float):
        if value < 0:
            raise ValueError("radius must be >= 0")
        self._radius = value

    @property
    def area(self) -> float:
        return 3.14159 * self._radius ** 2

c = Circle(5)
print(c.area)   # 78.53975(像属性一样访问)
c.radius = 10   # 触发 setter

四、迭代器与生成器(Think Python Ch19) ​

生成器是 Python 最强大的特性之一,它的核心思想是惰性求值——不是一次性计算所有结果存入内存,而是"按需生产"。这对处理大规模数据流至关重要。

理解生成器的关键:yield 关键字将普通函数变为"可暂停的函数"。每次调用 next() 时,生成器从上次 yield 处恢复执行,产生下一个值后再次暂停。这种"协程"思想后来发展为 async/await(第九章)。

与 Go 的对比:Go 用 channel + goroutine 实现生产者-消费者模式,Python 用生成器。两者都是"流式处理",但 Go 是并发的多 goroutine 通信,Python 生成器是单线程协作式——更简单但无法利用多核。

4.1 迭代器协议 ​

python
# for x in obj 等价于:
it = iter(obj)        # 调用 obj.__iter__() 返回迭代器
while True:
    try:
        x = next(it)  # 调用 it.__next__()
    except StopIteration:
        break

4.2 生成器(yield) ​

python
# 生成器函数:包含 yield → 返回生成器对象
def fibonacci(n: int):
    """生成前 n 个斐波那契数(惰性求值)"""
    a, b = 0, 1
    for _ in range(n):
        yield a
        a, b = b, a + b

# 生成器是单向的:只能产出值,不能接收值(除非用 send())

# 生成器表达式
squares = (x**2 for x in range(10**6))  # 不立刻分配内存
sum(x**2 for x in range(10**6))         # 直接传给函数

五、迭代器与生成器技巧(Cookbook Ch4) ​

5.1 手动消费迭代器 ​

python
items = [1, 2, 3]
it = iter(items)
next(it, None)          # 1(默认值 None 防止 StopIteration)
next(it, None)          # 2

5.2 itertools 模块 ​

python
from itertools import islice, chain, zip_longest, product, combinations

# 对迭代器切片
first_5 = list(islice(iterator, 5))

# 链式迭代
for x in chain([1, 2], "ab", {3, 4}):
    print(x)  # 1 2 a b 3 4

# 排列组合
list(product("AB", "12"))        # [('A','1'),('A','2'),('B','1'),('B','2')]
list(combinations("ABC", 2))     # [('A','B'),('A','C'),('B','C')]

六、文件 I/O 与数据序列化(Cookbook Ch5-6) ​

Python 的文件 I/O 设计体现了"简洁优于复杂"的哲学。with 语句自动管理资源生命周期(即使发生异常也会关闭文件),比 Java 的 try-with-resources 更早出现,比 Go 的 defer f.Close() 更安全(defer 只在函数返回时执行,with 在退出代码块时立即执行)。

什么是"Pythonic"的文件操作:用 for line in file 替代 while True: line = f.readline(),前者惰性逐行读取,不会把整个文件加载到内存。对大文件(GB 级别),这是唯一正确的方式。

6.1 文件读写的 Pythonic 方式 ​

python
# with 自动关闭文件
with open("data.txt", "r", encoding="utf-8") as f:
    for line in f:           # 惰性逐行读取,不占内存
        process(line)

# 一次读全部
with open("data.txt") as f:
    text = f.read()          # 适合小文件

# 写文件
with open("output.txt", "w") as f:
    f.write("hello\n")
    print("world", file=f)   # print 写入文件

6.2 JSON / CSV / Pickle ​

python
import json, csv

# JSON
data = json.loads('{"name": "Alice", "age": 30}')
json.dumps(data, indent=2, ensure_ascii=False)

# CSV
with open("data.csv") as f:
    reader = csv.DictReader(f)
    for row in reader:
        print(row["column_name"])

# Pickle(Python 专有,不能跨语言)
import pickle
pickle.dumps(obj)          # 序列化为 bytes
pickle.loads(data)         # 反序列化

七、函数进阶(Cookbook Ch7) ​

7.1 可调用对象 ​

python
# 自定义可调用对象(比闭包更清晰)
class Countdown:
    def __init__(self, start: int):
        self.start = start

    def __call__(self) -> bool:
        self.start -= 1
        return self.start > 0

cd = Countdown(5)
while cd():
    print("working...")  # 执行 4 次

7.2 偏函数与函数组合 ​

python
from functools import partial, reduce

# 偏函数:固定部分参数
parse_int_base2 = partial(int, base=2)
parse_int_base2("1010")   # 10

# reduce:累积计算
total = reduce(lambda a, b: a * b, [1, 2, 3, 4])  # 24

7.3 lambda vs 具名函数 ​

python
# lambda:仅用于简单表达式
sorted(users, key=lambda u: u.age)

# 如果逻辑变复杂 → 提取为函数
def get_fullname(user):
    return f"{user.last_name} {user.first_name}"
sorted(users, key=get_fullname)

八、类与对象进阶(Cookbook Ch8-9) ​

8.1 字符串表示 ​

python
class Point:
    def __repr__(self): return f"Point({self.x!r}, {self.y!r})"
    # !r 表示用 repr() 格式化,区别:
    # f"{s}"    → 调用 str(s)
    # f"{s!r}"  → 调用 repr(s)

8.2 创建管理型上下文管理器 ​

python
from contextlib import contextmanager

@contextmanager
def tempdir():
    """创建临时目录,自动清理"""
    import tempfile, shutil
    path = tempfile.mkdtemp()
    try:
        yield path
    finally:
        shutil.rmtree(path)

with tempdir() as d:
    open(f"{d}/file.txt", "w").write("hello")
# 自动清理

8.3 __slots__ 节约内存 ​

python
class Point:
    __slots__ = ["x", "y"]    # 禁止创建 __dict__,固定属性
    def __init__(self, x, y):
        self.x = x
        self.y = y

# 节省内存场景:创建数百万个简单对象时
# 常规类:每个实例有 __dict__(至少 64 字节 + 每个属性)
# __slots__:只有固定槽位(每个槽位 8 字节指针)

8.4 元编程:装饰器 ​

python
from functools import wraps

def log_call(func):
    """装饰器:记录每次调用"""
    @wraps(func)   # 保留原函数元信息(__name__, __doc__ 等)
    def wrapper(*args, **kwargs):
        print(f"calling {func.__name__}({args}, {kwargs})")
        return func(*args, **kwargs)
    return wrapper

@log_call
def add(a, b):
    """Return the sum of a and b."""
    return a + b

add(3, 5)  # calling add((3, 5), {})  → 8
help(add)  # "Return the sum of a and b."(wraps 保留了 docstring)

九、并发编程(Cookbook Ch12 + 深入) ​

9.1 GIL 原理 ​

GIL 是 CPython 的"阿喀琉斯之踵"——它让多线程在多核 CPU 上对 CPU 密集型任务完全失效。理解它不是"坏设计",而是引用计数式内存管理的历史代价。

mermaid
sequenceDiagram
    participant T1 as 线程1 (CPU密集)
    participant GIL as GIL
    participant T2 as 线程2 (CPU密集)
    participant OS as 操作系统

    OS->>T1: 调度到 CPU 核0
    T1->>GIL: 获取 GIL ✅
    T1->>T1: 执行 100 条字节码
    T1->>GIL: 释放 GIL
    Note over GIL,OS: 线程1被OS剥夺CPU
    OS->>T2: 调度到 CPU 核1
    T2->>GIL: 获取 GIL ✅
    T2->>T2: 执行 100 条字节码
    T2->>GIL: 释放 GIL
    Note over GIL,OS: ← 两个线程轮流执行,但绝不同时!
    Note over T1,T2: 多核CPU上,实际只有一个核在工作<br/>线程切换反而增加开销
GIL 关键事实:
  - CPython 解释器级别的互斥锁,每个 Python 进程只有一个
  - 同一时刻只有一个线程能执行 Python 字节码
  - 原因:CPython 的引用计数不是线程安全的,GIL 是保护它的"大锁"
  - 后果:
    · CPU 密集型多线程 → 无效(甚至比单线程慢,因为有切换开销)
    · I/O 密集型多线程 → 有效(I/O 等待时主动释放 GIL)
  - 绕过方式:多进程(multiprocessing)、C 扩展释放 GIL、用 PyPy/Jython(无 GIL)
  - Python 3.13 引入"Free-threaded CPython"(实验性,可禁用 GIL)
text
GIL 的底层实现:
  CPython 源码 (Python/ceval_gil.h):
    - 本质是一个非递归互斥锁 (PyMutex)
    - 每执行 100 条字节码 (sys.getswitchinterval() ≈ 5ms) 或遇到 I/O 操作时释放
    - 释放后立即重新尝试获取(其他线程有机会抢到)
    - 在 POSIX 上使用 pthread_mutex + pthread_cond 实现条件等待

  字节码执行循环 (ceval.c):
    while (1) {
        take_gil();           // 获取 GIL
        while (ticks-- > 0) {
            opcode = *next_instr++;
            switch (opcode) {
                case LOAD_FAST: ...
                case BINARY_ADD: ...
                // 每个 opcode 执行后检查是否需要释放 GIL
                if (ticks == 0) break; // 时间片用完 → 释放 GIL
            }
        }
        drop_gil();           // 释放 GIL → 其他线程有机会
    }

  引用计数为什么需要 GIL:
    PyObject->ob_refcnt 的 ++/-- 不是原子的(在大多数平台)。
    如果两个线程同时修改同一个对象的 ob_refcnt → 计数错乱
    → 内存泄漏或提前释放(UAF)。GIL 保证了每次只有一个人在改。

释放 GIL 的时机(由字节码解释器控制):
  - 执行了 100 条字节码(默认,可用 sys.setswitchinterval() 调整)
  - 调用 C 扩展中显式释放 GIL 的函数(Py_BEGIN_ALLOW_THREADS)
  - I/O 操作(read/write/connect/accept 等,CPython 在进入阻塞前释放 GIL)
  - time.sleep() — 主动让出

  为什么 I/O 多线程有效? → I/O 阻塞时 GIL 释放,其他线程获取 GIL 继续执行
  为什么 CPU 多线程无效? → 没有 I/O 释放点,只能等 100 条字节码配额用完

asyncio 事件循环的底层分析 ​

text
asyncio 本质上是在单线程中模拟并发——它不是真正的并行。

事件循环的核心循环 (简化版):
  while True:
    1. 检查就绪的 Future/Task(epoll_wait / select)
    2. 从就绪队列取一个 Task → 执行 Task.__step()
    3. Task 内部遇到 await → 挂起当前 Task → 注册回调
    4. 切换到下一个就绪 Task → 重复

await 到底做了什么?
  await asyncio.sleep(1) 等价于:
    1. 创建一个 TimerHandle(1秒后触发)
    2. 当前协程 yield 控制权 → 返回事件循环
    3. 1 秒后 TimerHandle 触发 → 事件循环恢复该协程

  与多线程的本质区别:
    - 协程切换: 用户态,无系统调用,~50ns
    - 线程切换: 内核态,保存/恢复上下文,~1-10μs
    - 但协程是协作式的: 一个协程死循环 → 整个事件循环卡死!
python
# asyncio 的微观示例 — 理解 await 挂起和恢复
import asyncio

async def probe():
    print(f"[0ms] probe start")
    await asyncio.sleep(0.01)  # ← 这里发生了什么?
    # 1. sleep(0.01) 返回一个 Future
    # 2. await 将当前协程"挂起"到该 Future
    # 3. 事件循环切换到其他就绪协程
    # 4. 0.01 秒后 Future 就绪 → 事件循环恢复此协程
    print(f"[10ms] probe resumed")
mermaid
sequenceDiagram
    participant Loop as 事件循环
    participant C1 as 协程1
    participant C2 as 协程2
    participant Timer as 定时器

    Loop->>C1: 恢复协程1
    C1->>C1: 执行字节码...
    C1->>Timer: await asyncio.sleep(0.01)
    Note over C1: 挂起, yield 给事件循环

    Loop->>C2: 恢复协程2
    C2->>C2: 执行字节码...

    Timer-->>Loop: 0.01s 后触发

    Loop->>C1: 协程1 恢复执行

并发模型决策树:

mermaid
flowchart TD
    Q["你的任务是?"]
    Q -->|"I/O 密集型<br/>(网络/文件/DB)"| IO["多线程<br/>ThreadPoolExecutor"]
    Q -->|"CPU 密集型<br/>(计算/图像)"| CPU["多进程<br/>ProcessPoolExecutor"]
    Q -->|"高并发 I/O<br/>(10000+连接)"| ASYNC["asyncio<br/>异步协程"]
    Q -->|"I/O + CPU 混合"| MIX["多进程 + 每进程内 asyncio"]

    IO --> IO2["✅ GIL 在 IO 时释放<br/>适合 Web 请求/爬虫"]
    CPU --> CPU2["✅ 绕过 GIL<br/>每个进程独立 GIL"]
    ASYNC --> ASYNC2["✅ 单线程协作式<br/>极高的 I/O 吞吐"]
    MIX --> MIX2["✅ 最佳并行<br/>各进程绑定 CPU+协程"]

    style IO2 fill:#4CAF50,color:#fff
    style CPU2 fill:#2196F3,color:#fff
    style ASYNC2 fill:#FF9800,color:#fff
    style MIX2 fill:#9C27B0,color:#fff

9.2 多线程(I/O 密集型) ​

python
from concurrent.futures import ThreadPoolExecutor, as_completed

def fetch_url(url: str) -> str:
    import urllib.request
    return urllib.request.urlopen(url).read().decode()

urls = ["https://httpbin.org/get"] * 10
with ThreadPoolExecutor(max_workers=5) as executor:
    futures = {executor.submit(fetch_url, url): url for url in urls}
    for future in as_completed(futures):
        result = future.result()

9.3 多进程(CPU 密集型) ​

python
from concurrent.futures import ProcessPoolExecutor

def compute_heavy(n: int) -> int:
    return sum(i * i for i in range(n))

with ProcessPoolExecutor(max_workers=4) as executor:
    results = list(executor.map(compute_heavy, [10**7] * 8))

9.4 asyncio(高并发 I/O) ​

asyncio 的核心是事件循环(Event Loop):一个单线程循环,管理着成百上千个协程,在它们之间来回切换。当协程遇到 await 时主动让出控制权,事件循环立即切换到另一个就绪的协程——这就是"协作式多任务"。

mermaid
sequenceDiagram
    participant EL as 事件循环 (单线程)
    participant C1 as 协程1 (fetch url1)
    participant C2 as 协程2 (fetch url2)
    participant IO as 网络 I/O

    EL->>C1: 执行到 await session.get()
    C1-->>EL: 让出控制权(yield)
    EL->>IO: 注册 fd1 到 epoll
    EL->>C2: 调度协程2
    C2->>C2: 执行到 await session.get()
    C2-->>EL: 让出控制权
    EL->>IO: 注册 fd2 到 epoll

    Note over EL: 无协程可调度,等待 I/O

    IO-->>EL: fd1 就绪!(数据到达)
    EL->>C1: 恢复协程1
    C1->>C1: 处理响应数据
    C1-->>EL: 完成

    IO-->>EL: fd2 就绪!
    EL->>C2: 恢复协程2
    C2->>C2: 处理响应数据
    C2-->>EL: 完成

    Note over EL: 全部协程完成,退出

关键对比:asyncio vs Go goroutine

asyncio (Python)goroutine (Go)
调度方式协作式(await 显式让出)抢占式(runtime 自动切换)
阻塞后果忘记 await = 阻塞整个事件循环goroutine 阻塞只挂起自身
CPU 利用单核多核
内存开销极小(协程栈 KB 级)小(goroutine 栈 2KB 起)
学习成本中等(理解 await/event loop)低(go 关键字即可)

核心原则:asyncio 中绝对不能调用阻塞函数(如 time.sleep()),必须使用 await asyncio.sleep()。否则整个事件循环会被卡住,所有协程全部停滞——这是新手最容易犯的错误。

python
import asyncio

async def fetch(session, url: str) -> str:
    async with session.get(url) as resp:
        return await resp.text()

async def main():
    import aiohttp
    async with aiohttp.ClientSession() as session:
        tasks = [fetch(session, f"https://httpbin.org/get?id={i}")
                 for i in range(100)]
        results = await asyncio.gather(*tasks)

asyncio.run(main())

9.5 GIL vs Go ​

维度Python (CPython)Go
并发模型线程 + GIL / asyncio 协程goroutine(抢占式)
并行多进程(有开销)多核并行(原生)
阻塞语义await 显式让出goroutine 自动切换
学习曲线需要理解 await/event loopgoroutine/channel 简洁

十、元编程(Cookbook Ch9) ​

10.1 类装饰器 ​

python
def add_repr(cls):
    """自动为类添加 __repr__ 方法"""
    def __repr__(self):
        attrs = ", ".join(f"{k}={v!r}" for k, v in self.__dict__.items()
                         if not k.startswith("_"))
        return f"{cls.__name__}({attrs})"
    cls.__repr__ = __repr__
    return cls

@add_repr
class Person:
    def __init__(self, name: str, age: int):
        self.name = name
        self.age = age

p = Person("Alice", 30)
print(p)  # Person(name='Alice', age=30)

10.2 元类 ​

python
# 元类 = 类的类(type 是默认元类)
# 场景:ORM(Django/SQLAlchemy)、接口强制执行

class ValidateFields(type):
    def __new__(cls, name, bases, attrs):
        for key, value in attrs.items():
            if key.startswith("__"):
                continue
            if not hasattr(value, "__annotations__"):
                raise TypeError(f"{key} must have type annotation")
        return super().__new__(cls, name, bases, attrs)

class Base(metaclass=ValidateFields):
    pass

十一、测试与调试(Think Python App B + Cookbook Ch14) ​

11.1 doctest ​

python
def add(a: int, b: int) -> int:
    """
    >>> add(1, 2)
    3
    >>> add(-1, 1)
    0
    """
    return a + b

# python -m doctest module.py

11.2 unittest / pytest ​

python
import pytest

def divide(a, b):
    if b == 0:
        raise ValueError("division by zero")
    return a / b

def test_divide():
    assert divide(10, 2) == 5.0
    with pytest.raises(ValueError):
        divide(10, 0)

@pytest.mark.parametrize("a,b,expected", [(10, 2, 5), (9, 3, 3)])
def test_parametrized(a, b, expected):
    assert divide(a, b) == expected

11.3 logging ​

python
import logging
logging.basicConfig(level=logging.INFO,
                    format="%(asctime)s [%(levelname)s] %(name)s: %(message)s")
logger = logging.getLogger(__name__)

logger.info("Starting process")
logger.warning("Low disk space")
logger.error("Connection failed", exc_info=True)  # 自动包含异常堆栈

十二、标准库精选(Cookbook 各章) ​

模块用途示例
pathlib面向对象路径操作Path("a/b").read_text()
subprocess子进程管理subprocess.run(["ls", "-l"], capture_output=True)
argparse命令行参数解析parser.add_argument("--verbose", action="store_true")
tempfile临时文件/目录tempfile.NamedTemporaryFile()
shutil高级文件操作shutil.copy2(src, dst)
datetime日期时间datetime.now(tz=timezone.utc)
collections高级容器deque, Counter, ChainMap
functools高阶函数lru_cache, partial, reduce
hashlib哈希摘要hashlib.sha256(data).hexdigest()
re正则表达式re.search(r"\d+", text)

argparse 完整示例 ​

python
import argparse

parser = argparse.ArgumentParser(description="Process some files")
parser.add_argument("files", nargs="+", help="input files")
parser.add_argument("-o", "--output", default="out.txt", help="output file")
parser.add_argument("-v", "--verbose", action="store_true", help="verbose output")

args = parser.parse_args(["-v", "-o", "result.txt", "a.txt", "b.txt"])
# args.files = ["a.txt", "b.txt"], args.output = "result.txt", args.verbose = True

十三、第三方库 ​

13.1 FastAPI(Web 框架) ​

python
from fastapi import FastAPI
from pydantic import BaseModel

app = FastAPI()

class Item(BaseModel):
    name: str
    price: float

@app.get("/items/{item_id}")
async def get_item(item_id: int) -> Item:
    return Item(name="Widget", price=19.99)

@app.post("/items")
async def create_item(item: Item) -> Item:
    # 自动校验、序列化
    return item

13.2 常用库速查 ​

库用途import
requestsHTTP 客户端requests.get(url)
aiohttp异步 HTTPasync with session.get(url)
pydantic数据校验BaseModel
pandas数据分析pd.read_csv("data.csv")
numpy数值计算np.array([1, 2, 3])
SQLAlchemyORM声明式映射
Celery分布式任务队列@app.task

十四、Python 内存管理深度 ​

Python 的内存管理是"引用计数 + 分代 GC"的双层机制——引用计数处理大部分对象的释放,分代 GC 兜底处理循环引用:

mermaid
flowchart TB
    subgraph RefCount["引用计数 (实时)"]
        RC1["每次赋值/传参: ob_refcnt++"]
        RC2["每次离开作用域: ob_refcnt--"]
        RC3["refcnt==0 → 立即释放 (tp_dealloc)"]
        RC4["✅ 延迟低, 内存释放及时"]
        RC5["❌ 循环引用无法释放!"]
    end

    subgraph GenGC["分代 GC (周期性)"]
        GC1["第0代: 新对象,频繁扫描"]
        GC2["第1代: 存活过的对象"]
        GC3["第2代: 长期存活对象,很少扫描"]
        GC4["用三色标记找到不可达的循环引用"]
    end

    RC5 -->|"触发 gc.collect()"| GC1

循环引用实战:

python
# 循环引用:内存泄漏的经典场景
class Node:
    def __init__(self, name):
        self.name = name
        self.ref = None

a = Node("A")
b = Node("B")
a.ref = b  # a → b
b.ref = a  # b → a  ← 循环引用!
# a 和 b 的 refcnt 永远 ≥ 1 → 引用计数不会释放
# 需要分代 GC 扫描才能找到并回收

CPython vs PyPy vs Cython — 性能对比 ​

维度CPython (3.12)PyPy (3.10)Cython
JIT 编译❌✅ 追踪 JIT❌ (AOT 编译为 C 扩展)
纯 Python 性能1x (基准)3-5xN/A
C 扩展兼容✅ 全兼容🟡 部分兼容(C API 慢)✅ 生成 C 扩展
GIL✅ 有 (3.13 可选无 GIL)❌ 无 GIL取决于 C 代码
内存占用中等高(JIT 编译缓存)低(接近 C)
启动速度✅ 极快❌ 慢(预热 JIT)✅ 快
适用通用 Python 代码纯 Python 计算密集型需要 C 速度的 Python

PyPy 不是万能加速器:它加速的是纯 Python 循环和计算(如 for 循环中的数值计算)。如果你的代码大量使用 C 扩展库(numpy/pandas)或 I/O 密集型——PyPy 不会带来明显提升。Cython 适合将某个 CPU 热点函数编译为 C 扩展。

十五、Python 与其他语言对比 ​

特性PythonGoC++Java
类型系统动态(鸭子类型)静态静态静态
并发模型GIL + asynciogoroutine线程 + 锁线程 + Executor
内存管理引用计数 + 分代 GCGC(并发标记-清除)RAII / 手动GC
编译解释 → 字节码编译 → 原生编译 → 原生编译 → 字节码
dict 实现开放寻址(有序)链地址法(无序)std::unordered_mapHashMap
list 扩容~1.125x<256→2x, ≥256→1.25x2x (GCC)1.5x
适合场景脚本/AI/Web后端微服务/基础设施系统/游戏/高频企业级后端

参考 ​

批注模式

💬 文章评论

暂无评论,来说点什么吧 👇

编程学习笔记