Skip to content

MongoDB ​

#数据库 · #MongoDB · #NoSQL · #文档数据库 · #分片 · #聚合 · #WiredTiger

MongoDB 是最流行的文档数据库。它以 BSON 格式存储数据,天然支持 JSON 风格查询,横向扩展能力强。理解其 WiredTiger 存储引擎、副本集和分片机制,是使用好 MongoDB 的关键。


核心概念 ​

MongoDBMySQL说明
DatabaseDatabase数据库
CollectionTable集合(表)
DocumentRow文档(行)
FieldColumn字段
_idPrimary Key主键(默认 ObjectId)
IndexIndex索引
Embedded DocJOIN内嵌文档替代关联

文档示例 ​

json
{
  "_id": ObjectId("507f1f77bcf86cd799439011"),
  "name": "高性能Go语言",
  "price": 59.00,
  "tags": ["go", "programming", "performance"],
  "author": {
    "name": "张三",
    "email": "zhangsan@example.com"
  },
  "reviews": [
    { "user": "alice", "rating": 5, "comment": "非常好" },
    { "user": "bob", "rating": 4 }
  ],
  "created_at": ISODate("2024-01-01T00:00:00Z")
}

CRUD 操作 ​

查询 ​

javascript
// 精确查询
db.products.find({ name: "高性能Go语言" })

// 范围查询
db.products.find({ price: { $gte: 30, $lt: 80 } })

// 数组查询
db.products.find({ tags: "go" })               // 数组包含
db.products.find({ tags: { $all: ["go", "performance"] } })  // 同时包含
db.products.find({ tags: { $size: 3 } })        // 数组长度

// 嵌套文档查询
db.products.find({ "author.name": "张三" })
db.products.find({ "reviews.rating": { $gte: 4 } })

// 投影(只返回部分字段)
db.products.find({}, { name: 1, price: 1, _id: 0 })

// 排序、分页
db.products.find().sort({ price: -1 }).skip(10).limit(5)

更新 ​

javascript
// 更新字段
db.products.updateOne(
    { _id: ObjectId("...") },
    { $set: { price: 49.99 } }
)

// 数组操作
db.products.updateOne(
    { _id: ObjectId("...") },
    { $push: { tags: "database" } }        // 添加元素
)

db.products.updateOne(
    { _id: ObjectId("...") },
    { $pull: { tags: "database" } }        // 删除元素
)

// upsert:不存在则插入
db.products.updateOne(
    { name: "新书" },
    { $set: { price: 99 } },
    { upsert: true }
)

聚合管道 ​

javascript
// 统计每个分类的书籍数量和均价
db.products.aggregate([
    // Stage 1: 筛选
    { $match: { price: { $gt: 0 } } },

    // Stage 2: 展开数组(如果 tags 是数组)
    { $unwind: "$tags" },

    // Stage 3: 分组
    { $group: {
        _id: "$tags",
        count: { $sum: 1 },
        avgPrice: { $avg: "$price" },
        minPrice: { $min: "$price" },
        maxPrice: { $max: "$price" }
    }},

    // Stage 4: 排序
    { $sort: { count: -1 } },

    // Stage 5: 限制
    { $limit: 10 }
])

常用聚合阶段 ​

阶段作用类似 SQL
$match过滤WHERE
$group分组GROUP BY
$sort排序ORDER BY
$limit限制行数LIMIT
$project投影/转换SELECT
$unwind展开数组UNNEST
$lookup跨集合关联LEFT JOIN
$addFields添加字段计算列
javascript
// $lookup 示例:关联查询
db.orders.aggregate([
    { $lookup: {
        from: "products",
        localField: "product_id",
        foreignField: "_id",
        as: "product"
    }}
])

索引 ​

javascript
// 单字段索引
db.products.createIndex({ name: 1 })  // 1=升序, -1=降序

// 复合索引(最左前缀匹配)
db.products.createIndex({ category: 1, price: -1 })

// 多键索引(数组字段自动)
db.products.createIndex({ tags: 1 })  // 数组字段自动为多键索引

// 文本索引
db.products.createIndex({ name: "text", description: "text" })
db.products.find({ $text: { $search: "go programming" } })

// 地理空间索引
db.places.createIndex({ location: "2dsphere" })
db.places.find({
    location: {
        $near: { $geometry: { type: "Point", coordinates: [116.4, 39.9] } }
    }
})

// 查看查询是否使用索引
db.products.find({ name: "xxx" }).explain("executionStats")

explain 关键指标 ​

executionStats:
  nReturned: 1            ← 返回文档数
  totalDocsExamined: 1    ← 扫描文档数(越小越好)
  totalKeysExamined: 1    ← 扫描索引键数
  executionTimeMillis: 0  ← 执行时间
  stage: "IXSCAN"         ← 索引扫描 ✅
  stage: "COLLSCAN"       ← 全表扫描 ❌

WiredTiger 存储引擎深度 ​

MongoDB 在 3.2 版本从 MMAPv1 切换到 WiredTiger,这不仅是"换一个存储引擎"——它改变了 MongoDB 的并发模型、压缩能力和崩溃恢复机制。

WiredTiger B-Tree vs 其他引擎 ​

WiredTiger 选择的是 B-Tree(实际上是 B+ Tree 的变体),而并非像 LevelDB/RocksDB 那样使用 LSM-Tree。这是一个权衡:

维度B-Tree (WiredTiger)LSM-Tree (RocksDB/LevelDB)
写入性能🟡 随机写需要多次寻道✅✅ 顺序写(追加写 + Compaction)
读取性能✅✅ O(log n) 一次查找🟡 可能需要查多层 + Bloom Filter
空间放大✅ 低(页内碎片可控)❌ 高(旧版本数据积压)
写放大✅ 低(就地更新)❌ 高(Compaction 重复写)
MVCC 实现页内多版本(Copy-on-Write 页分裂)天然多版本(SST 文件不可变)
压缩✅ snappy/zlib(页级压缩)✅ 块级压缩(更高效)
适合场景读写均衡、需要 MVCC写多读少、时序/日志数据

为什么 MongoDB 不选 LSM-Tree:MongoDB 的定位是"通用文档数据库",既要做 CRUD 查询、聚合管道,也要支撑事务——这些都需要较强的随机读能力。LSM-Tree 的读放大在高频点查场景下表现不佳。但如果你有大量写入+批量扫描的需求(如时序数据),应该考虑 ClickHouse(MergeTree)而非 MongoDB。

写入路径详解 ​

mermaid
flowchart TB
    Client["客户端写入"] --> Journal["Journal (WAL)<br/>磁盘持久化,顺序写<br/>用于崩溃恢复"]

    Client --> MemCache["WiredTiger Cache<br/>(内存中 B-Tree 页)"]

    MemCache -->|"Checkpoint<br/>每 60s 或 2GB WAL"| DataFiles["数据文件 (B-Tree)<br/>compressed (snappy/zlib)"]

    MemCache -->|"Eviction<br/>缓存满时淘汰脏页"| DataFiles

    Journal -.->|"崩溃恢复: 重放 Journal"| MemCache

MongoDB 的"写确认"三个等级:w:1(默认,Primary 确认即可)→ w:majority(多数节点确认)→ w:"majority", j:true(多数节点确认 + Journal 刷盘)。交易场景至少用 w:majority 才是真正可靠。

分片键选择策略 — 选错了比不分片更糟 ​

分片键决定了数据的物理分布——选错分片键比不分片还糟,因为热点分片会导致单节点过载:

mermaid
flowchart TD
    Q["你的查询模式?"]
    Q -->|"大部分查询按 user_id"| UserShard["✅ 哈希分片 {user_id: hashed}<br/>写入均匀, 单用户查询路由精准"]
    Q -->|"大部分查询按时间范围"| TimeShard["✅ 范围分片 {created_at: 1}<br/>时间范围查询高效, 但写入热点!"]
    Q -->|"混合: 按用户查 + 按时间查"| Hybrid["✅ 复合分片 {user_id: 1, created_at: 1}<br/>用户查询定点路由,时间过滤高效"]

    TimeShard --> HotSpot["⚠️ 热点问题: 新数据总写入最新分片<br/>解决: 用哈希复合分片<br/>{created_at: hashed, ...}"]

典型翻车案例:

javascript
// ❌ 糟糕的分片键
{ _id: 1 }               // 自增 ObjectId → 写入永远打到最后一个分片
{ status: 1 }            // 基数极低(只有几种状态) → 无法均匀分布

// ✅ 好的分片键
{ user_id: "hashed" }    // 高基数 + 哈希均匀
{ org_id: 1, _id: 1 }    // 按组织分区 + 文档排序

MongoDB vs MySQL vs PostgreSQL — 三数据库选型 ​

很多团队在"是否用 MongoDB"上反复纠结。本质不是"MongoDB 和 MySQL 谁更好",而是你的数据长什么样:

维度MongoDBMySQL (InnoDB)PostgreSQL
数据模型BSON 文档(Schema-less)关系表(强 Schema)关系表 + JSONB 混合
关联查询❌ $lookup 性能差(左连接)✅ JOIN 高效✅✅ JOIN + CTE + 窗口函数
Schema 变更✅✅ 零停机(字段随意加)❌ ALTER TABLE 锁表🟡 ALTER 支持并发但复杂
嵌套结构✅✅ 天然支持(内嵌文档)❌ 需拆表 + JOIN✅ JSONB 支持嵌套
事务 (ACID)✅ 4.0+ 多文档事务✅✅ 最成熟✅✅ 最成熟
水平扩展✅✅ 原生分片❌ 需中间件(ShardingSphere/Vitess)❌ 同上
文本搜索✅ 内置文本索引❌ FULLTEXT 弱✅✅ 原生全文搜索 (tsvector)
分析查询🟡 聚合管道(不如 SQL)🟡 基本✅✅ 窗口函数 + 分析函数
运维复杂度🟢 副本集/分片一键🟡 主从复制需手动🟡 类似 MySQL
生态系统🟡 较新但成长快✅✅ 最成熟✅ 非常成熟

选型速查:

你的数据长什么样?
├── 层级嵌套(订单含多个商品)→ MongoDB ✅
├── 严格关系(用户-角色-权限) → PostgreSQL/MySQL ✅
├── Schema 频繁变化(爬虫数据) → MongoDB ✅
├── 需要复杂报表/分析 → PostgreSQL ✅
├── 需要水平扩展(TB级数据) → MongoDB (分片) ✅
├── 需要全文搜索 → PostgreSQL ✅ 或 Elasticsearch
├── 需要强 ACID + JOIN → PostgreSQL >= MySQL > MongoDB

关键认知:MongoDB 的强项不在"替代 MySQL"——而是在 MySQL 做起来很痛苦的地方(Schemaless、嵌套文档、水平分片)。最佳实践是"各司其职":核心交易数据走 PostgreSQL,用户画像/日志/爬虫数据走 MongoDB,搜索引擎走 Elasticsearch。

事务 ​

javascript
// 多文档事务(MongoDB 4.0+)
const session = db.getMongo().startSession()
session.startTransaction()

try {
    session.getDatabase("shop").orders.insertOne({ ... })
    session.getDatabase("shop").inventory.updateOne(
        { product_id: 1 },
        { $inc: { stock: -1 } },
        { session }
    )
    session.commitTransaction()
} catch (e) {
    session.abortTransaction()
} finally {
    session.endSession()
}

副本集(Replica Set) ​

           ┌───────────┐
写入 ───→ │  Primary   │
           └─────┬─────┘
                 │ oplog 同步
      ┌──────────┼──────────┐
      ▼          ▼          ▼
┌──────────┐ ┌──────────┐ ┌──────────┐
│Secondary │ │Secondary │ │Secondary │
│  (data)  │ │  (data)  │ │ (arbiter)│
└──────────┘ └──────────┘ └──────────┘
                             无数据,仅投票

选举机制 ​

1. Primary 心跳超时(10s 无响应)
2. 有投票权的 Secondary 发起选举(Raft 协议)
3. 获得多数票的节点成为新 Primary
4. 客户端自动发现新 Primary(驱动层)

注意:至少需要 3 个节点(或 2 节点+1 arbiter)

分片(Sharding) ​

javascript
// 分片架构
//                 ┌──────────┐
//                 │  mongos  │ ← 路由(客户端连接)
//                 └────┬─────┘
//         ┌────────────┼────────────┐
//   ┌─────┴─────┐     │      ┌─────┴─────┐
//   │  Config   │     │      │  Config    │ ← 元数据
//   │  Server   │     │      │  Server    │
//   └───────────┘     │      └───────────┘
//                     │
//        ┌────────────┼────────────┐
//   ┌────┴────┐  ┌────┴────┐  ┌────┴────┐
//   │ Shard 1 │  │ Shard 2 │  │ Shard 3 │  ← 每个 Shard 是一个副本集
//   │  (0-500) │  │(500-1000)│  │(1000-...)│
//   └─────────┘  └─────────┘  └─────────┘

分片键选择 ​

类型分片键写入查询
范围分片{ _id: 1 }热点范围查询好
哈希分片{ _id: "hashed" }均匀范围查询差
复合分片{ user_id: 1, _id: 1 }按 user 分布用户查询好

常见问题 ​

1. 数据膨胀 ​

WiredTiger 的 MVCC 和删除标记导致磁盘空间增长
→ 定期 compact
db.runCommand({ compact: "collection_name" })

2. 慢查询 ​

javascript
// 开启慢查询日志
db.setProfilingLevel(1, { slowms: 100 })

// 查看慢查询
db.system.profile.find().sort({ ts: -1 }).limit(5)

3. 连接数 ​

bash
# 查看连接
db.serverStatus().connections
# { "current": 42, "available": 51158, "totalCreated": 100 }

与 MySQL 场景对比 ​

场景推荐
严格事务、复杂关联查询MySQL
Schema 灵活、嵌套文档MongoDB ✅
日志/事件存储MongoDB (TTL索引自动清理)
商品目录(属性可变)MongoDB ✅
财务/账务系统MySQL
实时分析、地理查询MongoDB ✅

参考 ​

批注模式

💬 文章评论

暂无评论,来说点什么吧 👇

编程学习笔记