第二阶段:经典神经网络
2.1 生物神经元 → 人工神经元
mermaid
graph LR
subgraph "生物神经元"
D1["树突<br/>接收信号"] --> BODY["细胞体<br/>整合信号"]
BODY --> A1["轴突<br/>输出信号"]
end
subgraph "人工神经元(感知机)"
X1["x₁"] --> S1["w₁"]
X2["x₂"] --> S2["w₂"]
X3["x₃"] --> S3["w₃"]
S1 --> SUM["∑ = Σwᵢxᵢ + b"]
S2 --> SUM
S3 --> SUM
SUM --> ACT["激活函数 σ"]
ACT --> OUT["输出 ŷ"]
end
style BODY fill:#3498db,color:#fff
style SUM fill:#e74c3c,color:#fff感知机 (Perceptron, 1958)
最简单的神经网络——只有一个神经元:
:输入特征 :权重(学习参数) :偏置 :激活函数(最早是阶跃函数)
感知机的致命缺陷:只能解决线性可分问题(1969 年 Minsky 证明了 XOR 问题无法用单层感知机解决),直接导致了第一次 AI 寒冬。
多层感知机 (MLP, 1986)
通过堆叠多个神经元层,MLP 解决了 XOR 问题:
mermaid
graph TD
subgraph "输入层"
X1["x₁"]
X2["x₂"]
end
subgraph "隐藏层1"
H1_1["h₁₁"]
H1_2["h₁₂"]
H1_3["h₁₃"]
end
subgraph "隐藏层2"
H2_1["h₂₁"]
H2_2["h₂₂"]
end
subgraph "输出层"
Y1["ŷ₁"]
end
X1 --> H1_1
X1 --> H1_2
X1 --> H1_3
X2 --> H1_1
X2 --> H1_2
X2 --> H1_3
H1_1 --> H2_1
H1_1 --> H2_2
H1_2 --> H2_1
H1_2 --> H2_2
H1_3 --> H2_1
H1_3 --> H2_2
H2_1 --> Y1
H2_2 --> Y1
style H1_1 fill:#3498db,color:#fff
style H1_2 fill:#3498db,color:#fff
style H1_3 fill:#3498db,color:#fff
style H2_1 fill:#9b59b6,color:#fff
style H2_2 fill:#9b59b6,color:#fff万能近似定理:只要隐藏层足够宽,一个单隐藏层的 MLP 就能以任意精度逼近任何连续函数。但要记住:这只是一个存在性定理,不告诉你需要多宽,也不保证学习算法能找到。
python
import torch.nn as nn
class MLP(nn.Module):
"""标准的多层感知机 — 深度学习的基础模块"""
def __init__(self, input_dim: int, hidden_dims: list, output_dim: int,
activation=nn.ReLU, dropout: float = 0.0):
super().__init__()
layers = []
dims = [input_dim] + hidden_dims
# 构建隐藏层
for i in range(len(hidden_dims)):
layers.append(nn.Linear(dims[i], dims[i+1]))
layers.append(activation())
if dropout > 0:
layers.append(nn.Dropout(dropout))
# 输出层(不加激活函数,由损失函数处理)
layers.append(nn.Linear(dims[-1], output_dim))
self.net = nn.Sequential(*layers)
def forward(self, x):
return self.net(x)
# 构建一个4层MLP
model = MLP(input_dim=784, hidden_dims=[512, 256, 128], output_dim=10, dropout=0.2)
print(f"网络结构:\n{model}")2.2 激活函数详解
激活函数是神经网络非线性的来源。没有激活函数,多层网络等价于单层(因为矩阵乘法满足结合律)。
mermaid
graph LR
subgraph "Sigmoid 族"
S1["Sigmoid<br/>σ(x) = 1/(1+e⁻ˣ)<br/>输出范围: (0,1)"]
S2["Tanh<br/>tanh(x) = (eˣ-e⁻ˣ)/(eˣ+e⁻ˣ)<br/>输出范围: (-1,1)"]
end
subgraph "ReLU 族"
R1["ReLU<br/>max(0,x)<br/>输出范围: [0,∞)"]
R2["Leaky ReLU<br/>max(0.01x,x)<br/>解决'死亡神经元'"]
R3["GELU<br/>x·Φ(x)<br/>Transformer 标配"]
end
subgraph "新一代"
G1["SiLU/Swish<br/>x·σ(x)<br/>LLaMA 等使用"]
G2["SwiGLU<br/>x·σ(βx)·W_g<br/>最新 LLM 标配"]
end
style R3 fill:#e74c3c,color:#fff
style G2 fill:#2ecc71,color:#fff每种激活函数的数学推导与使用场景
Sigmoid — 经典的"概率门"
python
import torch
import matplotlib.pyplot as plt
# 绘制五种激活函数
x = torch.linspace(-5, 5, 200)
activations = {
"Sigmoid": torch.sigmoid(x),
"Tanh": torch.tanh(x),
"ReLU": torch.relu(x),
"Leaky ReLU": torch.nn.functional.leaky_relu(x, 0.1),
"GELU": torch.nn.functional.gelu(x),
}
# 打印每个激活函数的特性
for name, y in activations.items():
print(f"{name:>12}: 范围 [{y.min():.2f}, {y.max():.2f}], "
f"x=0处导数约 {((y[101] - y[100]) / (x[101] - x[100])).item():.3f}")| 函数 | 公式 | 导数 | 优点 | 缺点 | 使用场景 |
|---|---|---|---|---|---|
| Sigmoid | 平滑、解释为概率 | 梯度消失(两端导数→0) | 输出层的二分类 | ||
| Tanh | 以0为中心 | 仍有梯度消失 | RNN 内部 | ||
| ReLU | 计算快、缓解梯度消失 | 死亡神经元(x<0 时梯度=0) | 通用隐藏层 | ||
| Leaky ReLU | 解决死亡神经元 | 多了超参数 | 替代 ReLU | ||
| GELU | 复杂 | 平滑、非单调 | 计算略贵 | Transformer 标配 | |
| SiLU (Swish) | 自门控、平滑 | 计算贵 | LLaMA 等新模型 | ||
| SwiGLU | — | 质量最优 | 参数量×2 | 最新 LLM 首选 |
关键理解:激活函数的演进方向是 "越来越平滑的非单调函数"。Sigmoid → ReLU → GELU → SwiGLU,每一步都提升了训练稳定性和模型质量。
2.3 损失函数 — "教会"模型什么是对的
损失函数衡量模型预测与真实答案之间的距离。训练 = 最小化损失函数。
均方误差 (MSE) — 回归任务
用于预测连续值:房价预测、温度预测。
交叉熵损失 — 分类任务
其中
python
import torch.nn as nn
# 回归损失
mse_loss = nn.MSELoss()
y_true = torch.tensor([2.5, 3.0, 4.0])
y_pred = torch.tensor([2.6, 2.8, 3.9])
print(f"MSE: {mse_loss(y_pred, y_true):.4f}")
# 分类损失
ce_loss = nn.CrossEntropyLoss()
logits = torch.tensor([[0.1, 0.2, 3.0], # 样本1:类别2概率最高
[2.0, 0.1, 0.05]]) # 样本2:类别0概率最高
labels = torch.tensor([2, 0]) # 真实标签
print(f"CrossEntropy: {ce_loss(logits, labels):.4f}")为什么分类用交叉熵而不是 MSE?
假设二分类,模型预测概率
2.4 反向传播 — 神经网络如何"学会"
mermaid
graph TD
subgraph "训练一轮的完整流程"
STEP1["1️⃣ 前向传播<br/>输入 x → 网络 → 预测 ŷ"] --> STEP2["2️⃣ 计算损失<br/>L = loss(ŷ, y_true)"]
STEP2 --> STEP3["3️⃣ 反向传播<br/>∂L/∂W 链式法则"]
STEP3 --> STEP4["4️⃣ 更新参数<br/>W ← W - η·∂L/∂W"]
STEP4 --> STEP1
end
style STEP3 fill:#e74c3c,color:#fff链式法则:反向传播的数学本质
反向传播是链式法则的工程实现。对于复合函数
每一步的"误差"从输出层逐层传回输入层。每一层都知道:我该往哪个方向调整参数,能让最终损失减小。
python
"""
手工实现反向传播 — 帮你理解 autograd 的原理
"""
import torch
# 简单例子: y = w², 求 dy/dw 在 w=3 处的值
w = torch.tensor([3.0], requires_grad=True)
y = w ** 2
y.backward() # 自动计算 dy/dw = 2w = 6
print(f"w=3 时 dy/dw = {w.grad.item()}") # 应输出 6.0
# 复杂例子: 模拟简单的两层网络
def manual_backprop():
"""手动演示反向传播的每一步"""
# 前向传播
x = torch.tensor([1.0, 2.0, 3.0]) # 输入
W1 = torch.tensor([[0.1, 0.2, 0.3],
[0.4, 0.5, 0.6]]) # 第一层权重
W2 = torch.tensor([0.7, 0.8]) # 第二层权重
y_true = torch.tensor([1.0]) # 真实值
# 前向传播
h = W1 @ x # (2,) 隐藏层输出
h_relu = torch.relu(h) # (2,) ReLU激活
y_pred = (W2 @ h_relu).reshape(1) # (1,) 最终输出
loss = (y_pred - y_true) ** 2 # MSE 损失
# 反向传播 (PyTorch 自动完成)
# 实际训练中只需要 loss.backward()
print(f"预测值: {y_pred.item():.4f}")
print(f"损失: {loss.item():.4f}")
print(f"\n损失对 W2 的梯度: ∂L/∂W2 = 2*(y_pred-y_true) * h_relu")
grad_W2 = 2 * (y_pred - y_true) * h_relu
print(f" = {grad_W2}")
print(f"\n损失对 W1 的梯度 (经过 ReLU 和 W2):")
# ∂L/∂W1 = ∂L/∂y_pred · ∂y_pred/∂h_relu · ∂h_relu/∂h · ∂h/∂W1
dL_dh_relu = 2 * (y_pred - y_true) * W2
dL_dh = dL_dh_relu * (h > 0).float() # ReLU 导数
grad_W1 = torch.outer(dL_dh, x)
print(f" = {grad_W1}")
# 取消注释以运行演示
# manual_backprop()梯度消失与梯度爆炸
mermaid
graph TD
subgraph "梯度消失 (Vanishing Gradient)"
V1["深层网络中<br/>Sigmoid σ' ∈ (0, 0.25)"] --> V2["链式法则连乘<br/>0.25¹⁰ ≈ 9.5×10⁻⁷"]
V2 --> V3["浅层参数几乎不更新<br/>训练停滞"]
end
subgraph "梯度爆炸 (Exploding Gradient)"
E1["权重初始化不当<br/>或 RNN 长序列"] --> E2["连乘结果 >> 1"]
E2 --> E3["梯度变成 NaN<br/>模型崩溃"]
end
subgraph "解决方案"
S1["✅ ReLU/GELU 激活"]
S2["✅ Batch/Layer Normalization"]
S3["✅ 残差连接 (Residual Connection)"]
S4["✅ 梯度裁剪 (Gradient Clipping)"]
end
style V3 fill:#e74c3c,color:#fff
style E3 fill:#e74c3c,color:#fff
style S1 fill:#2ecc71,color:#fff
style S2 fill:#2ecc71,color:#fff
style S3 fill:#2ecc71,color:#fff
style S4 fill:#2ecc71,color:#fff2.5 优化器 — 如何高效地沿着梯度下降
梯度下降 (GD)
最基础的优化算法——每一步沿着梯度反方向更新参数:
其中
三种梯度下降变体
| 变体 | 每次用的数据量 | 优点 | 缺点 |
|---|---|---|---|
| 批量 GD | 全部数据 | 稳定、收敛确定 | 太慢,内存放不下 |
| 随机 SGD | 1 个样本 | 快 | 震荡剧烈 |
| 小批量 SGD | 32-256 个样本 | 平衡速度与稳定性 | 需要调 batch size |
动量 (Momentum)
给梯度下降加上"惯性",像小球滚下山坡:
Adam — 当今最常用的优化器
Adam 结合了动量和自适应学习率:
python
import torch.optim as optim
# 常用优化器的 PyTorch 配置对比
optimizers = {
"SGD": optim.SGD(model.parameters(), lr=0.01),
"SGD+Momentum": optim.SGD(model.parameters(), lr=0.01, momentum=0.9),
"Adam": optim.Adam(model.parameters(), lr=0.001, betas=(0.9, 0.999)),
"AdamW": optim.AdamW(model.parameters(), lr=0.001, weight_decay=0.01),
}
# AdamW 是 Transformer 训练的首选
# weight_decay 与学习率解耦,比 Adam 的 L2 正则化效果更好学习率调度 (Learning Rate Scheduler)
mermaid
graph LR
A["Warmup<br/>lr 从 0 线性增加"] --> B["平稳阶段<br/>保持较高 lr"]
B --> C["衰减阶段<br/>cosine/线性衰减"]
style A fill:#f39c12,color:#fff
style B fill:#2ecc71,color:#fff
style C fill:#3498db,color:#fff典型配置(Transformer 训练):
python
from torch.optim.lr_scheduler import CosineAnnealingLR, LinearLR, SequentialLR
# Warmup + Cosine 衰减(Transformer 标准配置)
warmup = LinearLR(optimizer, start_factor=0.01, total_iters=1000)
cosine = CosineAnnealingLR(optimizer, T_max=9000)
scheduler = SequentialLR(optimizer, [warmup, cosine], milestones=[1000])2.6 卷积神经网络 (CNN) — 图像处理的基石
为什么需要 CNN?
MLP 处理图像的致命问题:一张 224×224×3 的图片,如果直接拉平送入全连接层,第一层的参数量为:
这不仅消耗巨大显存,还彻底丢失了空间结构信息。
CNN 通过两个核心思想解决这个问题:
- 局部感受野:每个神经元只看图像的一小块区域(如 3×3),而非整个图像
- 权重共享:同一个卷积核在图像上滑动,参数在所有位置共享
mermaid
graph TD
subgraph "MLP 的破坏性"
A1["224×224×3 = 150,528 像素"] --> A2["全连接: 1.5亿参数"]
A2 --> A3["❌ 空间信息丢失<br/>❌ 参数量爆炸"]
end
subgraph "CNN 的优雅"
B1["224×224×3"] --> B2["Conv3×3: 27参数<br/>在所有位置共享"]
B2 --> B3["✅ 保留空间结构<br/>✅ 参数高效"]
end
style A3 fill:#e74c3c,color:#fff
style B3 fill:#2ecc71,color:#fff卷积运算的数学定义
二维卷积(互相关)的定义:
其中
python
import torch
import torch.nn as nn
import torch.nn.functional as F
class SimpleCNN(nn.Module):
"""一个经典的 CNN 架构 — LeNet 风格的图像分类器"""
def __init__(self, num_classes: int = 10):
super().__init__()
# 第一卷积块:3 → 32 通道
self.conv1 = nn.Conv2d(3, 32, kernel_size=3, padding=1) # 输出: (B, 32, H, W)
self.bn1 = nn.BatchNorm2d(32)
self.pool1 = nn.MaxPool2d(2) # 输出: (B, 32, H/2, W/2)
# 第二卷积块:32 → 64 通道
self.conv2 = nn.Conv2d(32, 64, kernel_size=3, padding=1) # 输出: (B, 64, H/2, W/2)
self.bn2 = nn.BatchNorm2d(64)
self.pool2 = nn.MaxPool2d(2) # 输出: (B, 64, H/4, W/4)
# 第三卷积块:64 → 128 通道
self.conv3 = nn.Conv2d(64, 128, kernel_size=3, padding=1) # 输出: (B, 128, H/4, W/4)
self.bn3 = nn.BatchNorm2d(128)
self.pool3 = nn.AdaptiveAvgPool2d((4, 4)) # 输出: (B, 128, 4, 4)
# 分类头
self.fc1 = nn.Linear(128 * 4 * 4, 256)
self.dropout = nn.Dropout(0.5)
self.fc2 = nn.Linear(256, num_classes)
def forward(self, x: torch.Tensor) -> torch.Tensor:
# 卷积 → BN → ReLU → Pool
x = self.pool1(F.relu(self.bn1(self.conv1(x))))
x = self.pool2(F.relu(self.bn2(self.conv2(x))))
x = self.pool3(F.relu(self.bn3(self.conv3(x))))
# 展平 → 全连接
x = x.view(x.size(0), -1)
x = F.relu(self.fc1(x))
x = self.dropout(x)
x = self.fc2(x)
return x
# 测试
model = SimpleCNN(num_classes=10)
dummy = torch.randn(1, 3, 224, 224) # 模拟一张 224×224 图片
output = model(dummy)
print(f"输入形状: {dummy.shape}, 输出形状: {output.shape}")
print(f"总参数量: {sum(p.numel() for p in model.parameters()):,}")卷积操作的关键参数
| 参数 | 说明 | 典型值 | 效果 |
|---|---|---|---|
| kernel_size | 卷积核大小 | 3, 5, 7 | 越大感受野越大,计算量也越大 |
| padding | 边缘填充 | kernel_size // 2 | "same" padding 保持尺寸不变 |
| stride | 滑动步长 | 1(默认)/ 2 | stride=2 时尺寸减半,替代池化 |
| dilation | 空洞卷积 | 1(默认) | >1 时扩大感受野但不增加参数 |
输出尺寸计算公式
python
def calc_conv_output(H_in, kernel_size, stride=1, padding=0, dilation=1):
"""计算卷积输出尺寸"""
H_out = (H_in + 2 * padding - dilation * (kernel_size - 1) - 1) // stride + 1
return H_out
# 示例
print(f"224 → Conv3×3, padding=1, stride=1: {calc_conv_output(224, 3, padding=1)}")
print(f"112 → Conv3×3, padding=1, stride=2: {calc_conv_output(112, 3, stride=2, padding=1)}")池化 (Pooling) — 降采样
python
# 最大池化 vs 平均池化
x = torch.tensor([[[[1., 2., 3., 4.],
[5., 6., 7., 8.],
[9., 10., 11., 12.],
[13., 14., 15., 16.]]]])
# 2×2 最大池化
max_pool = F.max_pool2d(x, kernel_size=2, stride=2)
print(f"最大池化:\n{max_pool}") # 取每个 2×2 区域的最大值
# 2×2 平均池化
avg_pool = F.avg_pool2d(x, kernel_size=2, stride=2)
print(f"平均池化:\n{avg_pool}") # 取每个 2×2 区域的平均值
# 全局平均池化 (GAP) — 现代网络替代全连接层的常用方式
gap = F.adaptive_avg_pool2d(x, (1, 1)) # 任意尺寸 → 1×1
print(f"全局平均池化: {gap.shape}") # torch.Size([1, 1, 1, 1])经典 CNN 架构演进
mermaid
graph LR
LENET["LeNet-5 (1998)<br/>5层, 手写数字识别"] --> ALEX["AlexNet (2012)<br/>8层, ImageNet夺冠<br/>引入 ReLU+Dropout"]
ALEX --> VGG["VGG16/19 (2014)<br/>16层, 全3×3卷积"]
VGG --> INCEPT["Inception/GoogLeNet<br/>多尺度卷积"]
INCEPT --> RESNET["ResNet (2015)<br/>152层, 残差连接"]
RESNET --> EFF["EfficientNet<br/>神经架构搜索(NAS)"]
RESNET --> VIT["Vision Transformer<br/>CNN → Transformer"]
style ALEX fill:#e74c3c,color:#fff
style RESNET fill:#2ecc71,color:#fff残差块 (Residual Block) — 训练深层网络的关键
python
class ResidualBlock(nn.Module):
"""ResNet 核心模块 — 残差连接使训练 100+ 层成为可能"""
def __init__(self, in_channels: int, out_channels: int, stride: int = 1):
super().__init__()
self.conv1 = nn.Conv2d(in_channels, out_channels, 3, stride, 1)
self.bn1 = nn.BatchNorm2d(out_channels)
self.conv2 = nn.Conv2d(out_channels, out_channels, 3, 1, 1)
self.bn2 = nn.BatchNorm2d(out_channels)
# 如果维度不匹配,用 1×1 卷积对齐
self.shortcut = nn.Identity()
if stride != 1 or in_channels != out_channels:
self.shortcut = nn.Sequential(
nn.Conv2d(in_channels, out_channels, 1, stride),
nn.BatchNorm2d(out_channels),
)
def forward(self, x):
identity = self.shortcut(x)
out = F.relu(self.bn1(self.conv1(x)))
out = self.bn2(self.conv2(out))
out += identity # 关键:残差连接
return F.relu(out)2.7 正则化专题 — 防止过拟合的武器库
过拟合 = 模型在训练集上表现完美,但测试集上表现糟糕。本质是模型"记住了"训练数据的噪声而非"学会"了通用规律。
Dropout — 随机丢弃神经元
Dropout 在训练时以概率
python
# Dropout 的数学原理
# 训练时: 每个神经元以概率 p 被置零,剩余神经元输出乘以 1/(1-p) 保持期望不变
# 测试时: 所有神经元都工作,不进行任何丢弃
import torch.nn as nn
class MLPWithRegularization(nn.Module):
"""展示各种正则化技术的 MLP"""
def __init__(self, input_dim, hidden_dim, output_dim,
dropout_rate=0.5, use_bn=True):
super().__init__()
self.use_bn = use_bn
self.fc1 = nn.Linear(input_dim, hidden_dim)
self.bn1 = nn.BatchNorm1d(hidden_dim) if use_bn else nn.Identity()
self.dropout1 = nn.Dropout(dropout_rate) # Dropout
self.fc2 = nn.Linear(hidden_dim, hidden_dim)
self.bn2 = nn.BatchNorm1d(hidden_dim) if use_bn else nn.Identity()
self.dropout2 = nn.Dropout(dropout_rate)
self.fc3 = nn.Linear(hidden_dim, output_dim)
def forward(self, x):
x = self.dropout1(F.relu(self.bn1(self.fc1(x))))
x = self.dropout2(F.relu(self.bn2(self.fc2(x))))
x = self.fc3(x)
return xBatch Normalization — 批归一化
BatchNorm 对每个 mini-batch 的特征进行归一化,让每一层的输入分布保持稳定:
其中
为什么有效? BN 缓解了"内部协变量偏移"(Internal Covariate Shift),允许使用更大的学习率,减少对初始化的敏感性。
各正则化技术对比
| 技术 | 原理 | 适用场景 | 核心代码 |
|---|---|---|---|
| Dropout | 随机丢弃神经元 | 全连接层,Transformer | nn.Dropout(p=0.5) |
| BatchNorm | 按 batch 归一化 | CNN,大 batch | nn.BatchNorm2d(64) |
| LayerNorm | 按特征维度归一化 | Transformer,RNN | nn.LayerNorm(512) |
| Weight Decay | L2 惩罚权重 | 通用,AdamW | weight_decay=0.01 |
| Label Smoothing | 软化标签 | 分类任务 | label_smoothing=0.1 |
| Data Augmentation | 数据增广 | 图像/文本 | 随机裁剪、翻转、回译 |
| Early Stopping | 提前停止训练 | 通用 | 监控验证集 loss |
| Gradient Clipping | 梯度裁剪 | RNN,Transformer | clip_grad_norm_(max_norm=1.0) |
python
# 各种正则化在训练循环中的位置
def training_step_with_regularization(model, optimizer, batch, max_norm=1.0):
"""展示正则化技术的完整使用"""
inputs, targets = batch
# 前向传播(Dropout 在训练模式下自动生效)
model.train()
outputs = model(inputs)
loss = F.cross_entropy(outputs, targets, label_smoothing=0.1) # Label Smoothing
# 反向传播
optimizer.zero_grad()
loss.backward()
# 梯度裁剪
torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=max_norm)
# 参数更新(AdamW 的 weight_decay 在优化器内部处理)
optimizer.step()
return loss.item()权重初始化 — 从哪里开始很重要
| 方法 | 公式 | 适用激活函数 |
|---|---|---|
| Xavier/Glorot | Tanh, Sigmoid | |
| Kaiming/He | ReLU, Leaky ReLU | |
| Normal | Transformer (GPT 系列) |
python
def init_weights(module):
"""推荐的权重初始化函数"""
if isinstance(module, nn.Linear):
nn.init.kaiming_normal_(module.weight, mode='fan_out', nonlinearity='relu')
if module.bias is not None:
nn.init.zeros_(module.bias)
elif isinstance(module, nn.Conv2d):
nn.init.kaiming_normal_(module.weight, mode='fan_out', nonlinearity='relu')
if module.bias is not None:
nn.init.zeros_(module.bias)
elif isinstance(module, nn.Embedding):
nn.init.normal_(module.weight, mean=0, std=0.02) # Transformer 初始化
# 使用: model.apply(init_weights)2.8 完整的训练循环 — 从数据到模型
这是 d2l.ai 最核心的教学方式:让读者真正跑通一个训练流程。
python
"""
一个完整的深度学习训练循环
包含:数据加载 → 训练 → 验证 → 保存最佳模型 → 可视化
"""
import torch
import torch.nn as nn
import torch.optim as optim
from torch.utils.data import DataLoader, TensorDataset
import matplotlib.pyplot as plt
from tqdm import tqdm
import copy
class Trainer:
"""通用训练器 — 适用于任何 PyTorch 模型"""
def __init__(self, model, device='cuda' if torch.cuda.is_available() else 'cpu'):
self.model = model.to(device)
self.device = device
self.train_losses = []
self.val_losses = []
self.train_accs = []
self.val_accs = []
self.best_model_state = None
self.best_val_loss = float('inf')
def train_epoch(self, dataloader, optimizer, criterion):
"""一个训练 epoch"""
self.model.train()
total_loss, correct, total = 0.0, 0, 0
pbar = tqdm(dataloader, desc='Training', leave=False)
for inputs, targets in pbar:
inputs, targets = inputs.to(self.device), targets.to(self.device)
# 1. 前向传播
outputs = self.model(inputs)
loss = criterion(outputs, targets)
# 2. 反向传播
optimizer.zero_grad()
loss.backward()
# 3. 梯度裁剪(防止梯度爆炸)
torch.nn.utils.clip_grad_norm_(self.model.parameters(), max_norm=1.0)
# 4. 参数更新
optimizer.step()
# 5. 记录指标
total_loss += loss.item() * inputs.size(0)
pred = outputs.argmax(dim=1)
correct += (pred == targets).sum().item()
total += inputs.size(0)
pbar.set_postfix(loss=f'{loss.item():.4f}')
return total_loss / total, correct / total
@torch.no_grad()
def validate_epoch(self, dataloader, criterion):
"""一个验证 epoch"""
self.model.eval()
total_loss, correct, total = 0.0, 0, 0
for inputs, targets in tqdm(dataloader, desc='Validating', leave=False):
inputs, targets = inputs.to(self.device), targets.to(self.device)
outputs = self.model(inputs)
loss = criterion(outputs, targets)
total_loss += loss.item() * inputs.size(0)
pred = outputs.argmax(dim=1)
correct += (pred == targets).sum().item()
total += inputs.size(0)
return total_loss / total, correct / total
def fit(self, train_loader, val_loader, epochs=10, lr=0.001,
weight_decay=1e-4, lr_scheduler=None, patience=5):
"""完整的训练流程"""
optimizer = optim.AdamW(self.model.parameters(), lr=lr, weight_decay=weight_decay)
criterion = nn.CrossEntropyLoss(label_smoothing=0.1)
patience_counter = 0
print(f"🚀 开始训练 (设备: {self.device}, 模型参数: {sum(p.numel() for p in self.model.parameters()):,})")
print(f"{'Epoch':<8} {'Train Loss':<13} {'Train Acc':<11} {'Val Loss':<13} {'Val Acc':<11} {'LR':<10} {'Best':<6}")
print("-" * 72)
for epoch in range(1, epochs + 1):
# 训练
train_loss, train_acc = self.train_epoch(train_loader, optimizer, criterion)
self.train_losses.append(train_loss)
self.train_accs.append(train_acc)
# 验证
val_loss, val_acc = self.validate_epoch(val_loader, criterion)
self.val_losses.append(val_loss)
self.val_accs.append(val_acc)
# 学习率调度
current_lr = optimizer.param_groups[0]['lr']
if lr_scheduler:
lr_scheduler.step()
# 保存最佳模型
is_best = val_loss < self.best_val_loss
if is_best:
self.best_val_loss = val_loss
self.best_model_state = copy.deepcopy(self.model.state_dict())
patience_counter = 0
else:
patience_counter += 1
# 打印进度
best_mark = '✅' if is_best else ''
print(f"{epoch:<8} {train_loss:<13.4f} {train_acc:<11.4f} "
f"{val_loss:<13.4f} {val_acc:<11.4f} {current_lr:<10.2e} {best_mark:<6}")
# 早停
if patience_counter >= patience:
print(f"\n⏹️ 验证集 loss 连续 {patience} 轮未改善,提前停止训练")
break
# 恢复最佳模型
if self.best_model_state:
self.model.load_state_dict(self.best_model_state)
print(f"\n✅ 训练完成!最佳验证 loss: {self.best_val_loss:.4f}")
return self.model
def plot_metrics(self):
"""绘制训练曲线"""
fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(12, 4))
# Loss 曲线
ax1.plot(self.train_losses, label='Train Loss', color='#3498db')
ax1.plot(self.val_losses, label='Val Loss', color='#e74c3c')
ax1.set_xlabel('Epoch')
ax1.set_ylabel('Loss')
ax1.set_title('损失曲线')
ax1.legend()
ax1.grid(True, alpha=0.3)
# Accuracy 曲线
ax2.plot(self.train_accs, label='Train Acc', color='#3498db')
ax2.plot(self.val_accs, label='Val Acc', color='#2ecc71')
ax2.set_xlabel('Epoch')
ax2.set_ylabel('Accuracy')
ax2.set_title('准确率曲线')
ax2.legend()
ax2.grid(True, alpha=0.3)
plt.tight_layout()
plt.savefig('training_curves.png', dpi=150)
# plt.show()
# ========== 完整使用示例 ==========
def run_training_demo():
"""从数据生成到模型训练的完整示例"""
# 1. 生成模拟数据
print("📦 1. 生成数据...")
torch.manual_seed(42)
n_samples = 10000
n_features = 100
n_classes = 10
# 生成特征和标签
X = torch.randn(n_samples, n_features)
true_W = torch.randn(n_features, n_classes)
y = (X @ true_W).argmax(dim=1)
# 划分数据集
n_train = int(0.8 * n_samples)
train_X, train_y = X[:n_train], y[:n_train]
val_X, val_y = X[n_train:], y[n_train:]
train_dataset = TensorDataset(train_X, train_y)
val_dataset = TensorDataset(val_X, val_y)
train_loader = DataLoader(train_dataset, batch_size=64, shuffle=True)
val_loader = DataLoader(val_dataset, batch_size=128, shuffle=False)
# 2. 创建模型
print("🧠 2. 创建模型...")
model = MLP(input_dim=n_features, hidden_dims=[256, 128, 64],
output_dim=n_classes, dropout=0.3)
# 3. 设置学习率调度
optimizer_pre = optim.AdamW(model.parameters(), lr=0.001)
scheduler = optim.lr_scheduler.CosineAnnealingLR(optimizer_pre, T_max=20)
# 4. 训练
print("🚀 3. 开始训练...")
trainer = Trainer(model)
trainer.fit(
train_loader, val_loader,
epochs=20, lr=0.001, weight_decay=1e-4,
lr_scheduler=scheduler, patience=5,
)
# 5. 可视化
trainer.plot_metrics()
print("📊 训练曲线已保存为 training_curves.png")
# run_training_demo()训练循环的核心要点
mermaid
graph TD
subgraph "每个 epoch 的执行顺序"
S1["model.train()<br/>开启训练模式<br/>(Dropout/BatchNorm 生效)"] --> S2["for batch in dataloader:"]
S2 --> S3["outputs = model(inputs)<br/>前向传播"]
S3 --> S4["loss = criterion(outputs, targets)<br/>计算损失"]
S4 --> S5["optimizer.zero_grad()<br/>清零梯度"]
S5 --> S6["loss.backward()<br/>反向传播"]
S6 --> S7["clip_grad_norm_()<br/>梯度裁剪(可选)"]
S7 --> S8["optimizer.step()<br/>参数更新"]
S8 --> S2
S2 -->|epoch 结束| S9["model.eval()<br/>验证模式<br/>(Dropout 关闭)"]
S9 --> S10["scheduler.step()<br/>调整学习率"]
S10 --> S11["保存最佳模型<br/>(如果 val_loss 降低)"]
end
style S3 fill:#3498db,color:#fff
style S6 fill:#e74c3c,color:#fff
style S8 fill:#2ecc71,color:#fff
style S11 fill:#f39c12,color:#fff| 陷阱 | 说明 | 正确做法 |
|---|---|---|
忘记 zero_grad() | 梯度会累积而非替换 | 每次 backward() 前调用 |
| 验证时不关 Dropout | 输出不稳定,指标不准 | 验证前必须 model.eval() |
忘记 .to(device) | CPU/GPU 不一致报错 | .to(device) 统一 |
DataLoader 不 shuffle | 验证集不需要,训练集必须 | 训练: shuffle=True,验证: False |
| 梯度裁剪缺失 | RNN/深层网络容易梯度爆炸 | clip_grad_norm_(max_norm=1.0) |
🧭 学习导航
← 上一阶段:math-programming | 返回总览 | 下一阶段:nlp-sequence-models →
登录后即可发表评论 👇