MongoDB
#数据库 · #MongoDB · #NoSQL · #文档数据库 · #分片 · #聚合 · #WiredTiger
MongoDB 是最流行的文档数据库。它以 BSON 格式存储数据,天然支持 JSON 风格查询,横向扩展能力强。理解其 WiredTiger 存储引擎、副本集和分片机制,是使用好 MongoDB 的关键。
核心概念
| MongoDB | MySQL | 说明 |
|---|---|---|
| Database | Database | 数据库 |
| Collection | Table | 集合(表) |
| Document | Row | 文档(行) |
| Field | Column | 字段 |
_id | Primary Key | 主键(默认 ObjectId) |
| Index | Index | 索引 |
| Embedded Doc | JOIN | 内嵌文档替代关联 |
文档示例
json
{
"_id": ObjectId("507f1f77bcf86cd799439011"),
"name": "高性能Go语言",
"price": 59.00,
"tags": ["go", "programming", "performance"],
"author": {
"name": "张三",
"email": "zhangsan@example.com"
},
"reviews": [
{ "user": "alice", "rating": 5, "comment": "非常好" },
{ "user": "bob", "rating": 4 }
],
"created_at": ISODate("2024-01-01T00:00:00Z")
}CRUD 操作
查询
javascript
// 精确查询
db.products.find({ name: "高性能Go语言" })
// 范围查询
db.products.find({ price: { $gte: 30, $lt: 80 } })
// 数组查询
db.products.find({ tags: "go" }) // 数组包含
db.products.find({ tags: { $all: ["go", "performance"] } }) // 同时包含
db.products.find({ tags: { $size: 3 } }) // 数组长度
// 嵌套文档查询
db.products.find({ "author.name": "张三" })
db.products.find({ "reviews.rating": { $gte: 4 } })
// 投影(只返回部分字段)
db.products.find({}, { name: 1, price: 1, _id: 0 })
// 排序、分页
db.products.find().sort({ price: -1 }).skip(10).limit(5)更新
javascript
// 更新字段
db.products.updateOne(
{ _id: ObjectId("...") },
{ $set: { price: 49.99 } }
)
// 数组操作
db.products.updateOne(
{ _id: ObjectId("...") },
{ $push: { tags: "database" } } // 添加元素
)
db.products.updateOne(
{ _id: ObjectId("...") },
{ $pull: { tags: "database" } } // 删除元素
)
// upsert:不存在则插入
db.products.updateOne(
{ name: "新书" },
{ $set: { price: 99 } },
{ upsert: true }
)聚合管道
javascript
// 统计每个分类的书籍数量和均价
db.products.aggregate([
// Stage 1: 筛选
{ $match: { price: { $gt: 0 } } },
// Stage 2: 展开数组(如果 tags 是数组)
{ $unwind: "$tags" },
// Stage 3: 分组
{ $group: {
_id: "$tags",
count: { $sum: 1 },
avgPrice: { $avg: "$price" },
minPrice: { $min: "$price" },
maxPrice: { $max: "$price" }
}},
// Stage 4: 排序
{ $sort: { count: -1 } },
// Stage 5: 限制
{ $limit: 10 }
])常用聚合阶段
| 阶段 | 作用 | 类似 SQL |
|---|---|---|
$match | 过滤 | WHERE |
$group | 分组 | GROUP BY |
$sort | 排序 | ORDER BY |
$limit | 限制行数 | LIMIT |
$project | 投影/转换 | SELECT |
$unwind | 展开数组 | UNNEST |
$lookup | 跨集合关联 | LEFT JOIN |
$addFields | 添加字段 | 计算列 |
javascript
// $lookup 示例:关联查询
db.orders.aggregate([
{ $lookup: {
from: "products",
localField: "product_id",
foreignField: "_id",
as: "product"
}}
])索引
javascript
// 单字段索引
db.products.createIndex({ name: 1 }) // 1=升序, -1=降序
// 复合索引(最左前缀匹配)
db.products.createIndex({ category: 1, price: -1 })
// 多键索引(数组字段自动)
db.products.createIndex({ tags: 1 }) // 数组字段自动为多键索引
// 文本索引
db.products.createIndex({ name: "text", description: "text" })
db.products.find({ $text: { $search: "go programming" } })
// 地理空间索引
db.places.createIndex({ location: "2dsphere" })
db.places.find({
location: {
$near: { $geometry: { type: "Point", coordinates: [116.4, 39.9] } }
}
})
// 查看查询是否使用索引
db.products.find({ name: "xxx" }).explain("executionStats")explain 关键指标
executionStats:
nReturned: 1 ← 返回文档数
totalDocsExamined: 1 ← 扫描文档数(越小越好)
totalKeysExamined: 1 ← 扫描索引键数
executionTimeMillis: 0 ← 执行时间
stage: "IXSCAN" ← 索引扫描 ✅
stage: "COLLSCAN" ← 全表扫描 ❌WiredTiger 存储引擎深度
MongoDB 在 3.2 版本从 MMAPv1 切换到 WiredTiger,这不仅是"换一个存储引擎"——它改变了 MongoDB 的并发模型、压缩能力和崩溃恢复机制。
WiredTiger B-Tree vs 其他引擎
WiredTiger 选择的是 B-Tree(实际上是 B+ Tree 的变体),而并非像 LevelDB/RocksDB 那样使用 LSM-Tree。这是一个权衡:
| 维度 | B-Tree (WiredTiger) | LSM-Tree (RocksDB/LevelDB) |
|---|---|---|
| 写入性能 | 🟡 随机写需要多次寻道 | ✅✅ 顺序写(追加写 + Compaction) |
| 读取性能 | ✅✅ O(log n) 一次查找 | 🟡 可能需要查多层 + Bloom Filter |
| 空间放大 | ✅ 低(页内碎片可控) | ❌ 高(旧版本数据积压) |
| 写放大 | ✅ 低(就地更新) | ❌ 高(Compaction 重复写) |
| MVCC 实现 | 页内多版本(Copy-on-Write 页分裂) | 天然多版本(SST 文件不可变) |
| 压缩 | ✅ snappy/zlib(页级压缩) | ✅ 块级压缩(更高效) |
| 适合场景 | 读写均衡、需要 MVCC | 写多读少、时序/日志数据 |
为什么 MongoDB 不选 LSM-Tree:MongoDB 的定位是"通用文档数据库",既要做 CRUD 查询、聚合管道,也要支撑事务——这些都需要较强的随机读能力。LSM-Tree 的读放大在高频点查场景下表现不佳。但如果你有大量写入+批量扫描的需求(如时序数据),应该考虑 ClickHouse(MergeTree)而非 MongoDB。
写入路径详解
mermaid
flowchart TB
Client["客户端写入"] --> Journal["Journal (WAL)<br/>磁盘持久化,顺序写<br/>用于崩溃恢复"]
Client --> MemCache["WiredTiger Cache<br/>(内存中 B-Tree 页)"]
MemCache -->|"Checkpoint<br/>每 60s 或 2GB WAL"| DataFiles["数据文件 (B-Tree)<br/>compressed (snappy/zlib)"]
MemCache -->|"Eviction<br/>缓存满时淘汰脏页"| DataFiles
Journal -.->|"崩溃恢复: 重放 Journal"| MemCacheMongoDB 的"写确认"三个等级:
w:1(默认,Primary 确认即可)→w:majority(多数节点确认)→w:"majority", j:true(多数节点确认 + Journal 刷盘)。交易场景至少用w:majority才是真正可靠。
分片键选择策略 — 选错了比不分片更糟
分片键决定了数据的物理分布——选错分片键比不分片还糟,因为热点分片会导致单节点过载:
mermaid
flowchart TD
Q["你的查询模式?"]
Q -->|"大部分查询按 user_id"| UserShard["✅ 哈希分片 {user_id: hashed}<br/>写入均匀, 单用户查询路由精准"]
Q -->|"大部分查询按时间范围"| TimeShard["✅ 范围分片 {created_at: 1}<br/>时间范围查询高效, 但写入热点!"]
Q -->|"混合: 按用户查 + 按时间查"| Hybrid["✅ 复合分片 {user_id: 1, created_at: 1}<br/>用户查询定点路由,时间过滤高效"]
TimeShard --> HotSpot["⚠️ 热点问题: 新数据总写入最新分片<br/>解决: 用哈希复合分片<br/>{created_at: hashed, ...}"]典型翻车案例:
javascript
// ❌ 糟糕的分片键
{ _id: 1 } // 自增 ObjectId → 写入永远打到最后一个分片
{ status: 1 } // 基数极低(只有几种状态) → 无法均匀分布
// ✅ 好的分片键
{ user_id: "hashed" } // 高基数 + 哈希均匀
{ org_id: 1, _id: 1 } // 按组织分区 + 文档排序MongoDB vs MySQL vs PostgreSQL — 三数据库选型
很多团队在"是否用 MongoDB"上反复纠结。本质不是"MongoDB 和 MySQL 谁更好",而是你的数据长什么样:
| 维度 | MongoDB | MySQL (InnoDB) | PostgreSQL |
|---|---|---|---|
| 数据模型 | BSON 文档(Schema-less) | 关系表(强 Schema) | 关系表 + JSONB 混合 |
| 关联查询 | ❌ $lookup 性能差(左连接) | ✅ JOIN 高效 | ✅✅ JOIN + CTE + 窗口函数 |
| Schema 变更 | ✅✅ 零停机(字段随意加) | ❌ ALTER TABLE 锁表 | 🟡 ALTER 支持并发但复杂 |
| 嵌套结构 | ✅✅ 天然支持(内嵌文档) | ❌ 需拆表 + JOIN | ✅ JSONB 支持嵌套 |
| 事务 (ACID) | ✅ 4.0+ 多文档事务 | ✅✅ 最成熟 | ✅✅ 最成熟 |
| 水平扩展 | ✅✅ 原生分片 | ❌ 需中间件(ShardingSphere/Vitess) | ❌ 同上 |
| 文本搜索 | ✅ 内置文本索引 | ❌ FULLTEXT 弱 | ✅✅ 原生全文搜索 (tsvector) |
| 分析查询 | 🟡 聚合管道(不如 SQL) | 🟡 基本 | ✅✅ 窗口函数 + 分析函数 |
| 运维复杂度 | 🟢 副本集/分片一键 | 🟡 主从复制需手动 | 🟡 类似 MySQL |
| 生态系统 | 🟡 较新但成长快 | ✅✅ 最成熟 | ✅ 非常成熟 |
选型速查:
你的数据长什么样?
├── 层级嵌套(订单含多个商品)→ MongoDB ✅
├── 严格关系(用户-角色-权限) → PostgreSQL/MySQL ✅
├── Schema 频繁变化(爬虫数据) → MongoDB ✅
├── 需要复杂报表/分析 → PostgreSQL ✅
├── 需要水平扩展(TB级数据) → MongoDB (分片) ✅
├── 需要全文搜索 → PostgreSQL ✅ 或 Elasticsearch
├── 需要强 ACID + JOIN → PostgreSQL >= MySQL > MongoDB关键认知:MongoDB 的强项不在"替代 MySQL"——而是在 MySQL 做起来很痛苦的地方(Schemaless、嵌套文档、水平分片)。最佳实践是"各司其职":核心交易数据走 PostgreSQL,用户画像/日志/爬虫数据走 MongoDB,搜索引擎走 Elasticsearch。
事务
javascript
// 多文档事务(MongoDB 4.0+)
const session = db.getMongo().startSession()
session.startTransaction()
try {
session.getDatabase("shop").orders.insertOne({ ... })
session.getDatabase("shop").inventory.updateOne(
{ product_id: 1 },
{ $inc: { stock: -1 } },
{ session }
)
session.commitTransaction()
} catch (e) {
session.abortTransaction()
} finally {
session.endSession()
}副本集(Replica Set)
┌───────────┐
写入 ───→ │ Primary │
└─────┬─────┘
│ oplog 同步
┌──────────┼──────────┐
▼ ▼ ▼
┌──────────┐ ┌──────────┐ ┌──────────┐
│Secondary │ │Secondary │ │Secondary │
│ (data) │ │ (data) │ │ (arbiter)│
└──────────┘ └──────────┘ └──────────┘
无数据,仅投票选举机制
1. Primary 心跳超时(10s 无响应)
2. 有投票权的 Secondary 发起选举(Raft 协议)
3. 获得多数票的节点成为新 Primary
4. 客户端自动发现新 Primary(驱动层)
注意:至少需要 3 个节点(或 2 节点+1 arbiter)分片(Sharding)
javascript
// 分片架构
// ┌──────────┐
// │ mongos │ ← 路由(客户端连接)
// └────┬─────┘
// ┌────────────┼────────────┐
// ┌─────┴─────┐ │ ┌─────┴─────┐
// │ Config │ │ │ Config │ ← 元数据
// │ Server │ │ │ Server │
// └───────────┘ │ └───────────┘
// │
// ┌────────────┼────────────┐
// ┌────┴────┐ ┌────┴────┐ ┌────┴────┐
// │ Shard 1 │ │ Shard 2 │ │ Shard 3 │ ← 每个 Shard 是一个副本集
// │ (0-500) │ │(500-1000)│ │(1000-...)│
// └─────────┘ └─────────┘ └─────────┘分片键选择
| 类型 | 分片键 | 写入 | 查询 |
|---|---|---|---|
| 范围分片 | { _id: 1 } | 热点 | 范围查询好 |
| 哈希分片 | { _id: "hashed" } | 均匀 | 范围查询差 |
| 复合分片 | { user_id: 1, _id: 1 } | 按 user 分布 | 用户查询好 |
常见问题
1. 数据膨胀
WiredTiger 的 MVCC 和删除标记导致磁盘空间增长
→ 定期 compact
db.runCommand({ compact: "collection_name" })2. 慢查询
javascript
// 开启慢查询日志
db.setProfilingLevel(1, { slowms: 100 })
// 查看慢查询
db.system.profile.find().sort({ ts: -1 }).limit(5)3. 连接数
bash
# 查看连接
db.serverStatus().connections
# { "current": 42, "available": 51158, "totalCreated": 100 }与 MySQL 场景对比
| 场景 | 推荐 |
|---|---|
| 严格事务、复杂关联查询 | MySQL |
| Schema 灵活、嵌套文档 | MongoDB ✅ |
| 日志/事件存储 | MongoDB (TTL索引自动清理) |
| 商品目录(属性可变) | MongoDB ✅ |
| 财务/账务系统 | MySQL |
| 实时分析、地理查询 | MongoDB ✅ |
登录后即可发表评论 👇