Skip to content
第 11 章 ⏱ 14 分钟阅读

第 11 章:MongoDB 进阶 ​

学习目标 ​

  • 掌握聚合管道($match、$group、$lookup)
  • 设计合理的索引(单字段、复合、覆盖)
  • 用副本集保证高可用
  • 了解分片集群的水平扩展

一、聚合管道(Aggregation Pipeline) ​

聚合管道 = 一系列 stage,每个 stage 接收文档流,转换后传给下一个 stage。

javascript
db.orders.aggregate([
  { $match:   { status: "PAID" } },           // 过滤
  { $group:   { _id: "$userId", total: { $sum: "$amount" } } },  // 按用户分组求和
  { $sort:    { total: -1 } },                // 排序
  { $limit:   10 }                             // 取前 10
])

效果:每个付费用户的总消费额 Top 10。

二、常用 Stage ​

2.1 $match 过滤 ​

javascript
{ $match: { status: "PAID", createdAt: { $gte: ISODate("2024-01-01") } } }

⚠️ 坑 1:$match 尽量放在管道最前面,能利用索引,后面再 group。

2.2 $group 分组 ​

javascript
{
  $group: {
    _id: "$category",                  // 分组键
    count: { $sum: 1 },                // 计数
    avgPrice: { $avg: "$price" },      // 平均
    maxPrice: { $max: "$price" },
    items: { $push: "$name" }          // 收集到数组
  }
}

⚠️ 坑 2:$push 没限制会把所有文档字段塞进去,可能爆内存。生产用 $push + $slice 截断。

2.3 $project 投影 ​

javascript
{
  $project: {
    name: 1,                           // 1 = 包含
    price: 1,
    isExpensive: { $gt: ["$price", 100] }
  }
}

2.4 $lookup JOIN ​

javascript
{
  $lookup: {
    from:         "users",
    localField:   "userId",
    foreignField: "_id",
    as:           "userInfo"
  }
}

效果:把 order 的 userId 关联到 users 的 _id,结果存在 userInfo 数组里(永远数组)。

⚠️ 坑 3:$lookup 性能比 MySQL JOIN 差很多,不要用在高频查询,先考虑反范式设计。

2.5 $unwind 拆数组 ​

javascript
{ $unwind: "$items" }     // 把 items: [a, b] 拆成两条:items: a 和 items: b

三、聚合管道实战 ​

需求:每个用户的订单数、总金额、最近订单时间。

javascript
db.orders.aggregate([
  { $match: { status: "PAID" } },
  {
    $group: {
      _id: "$userId",
      orderCount: { $sum: 1 },
      totalAmount: { $sum: "$amount" },
      lastOrderAt: { $max: "$createdAt" }
    }
  },
  { $sort: { totalAmount: -1 } },
  { $limit: 20 }
])

四、Java 聚合 ​

java
import static org.springframework.data.mongodb.core.aggregation.Aggregation.*;

Aggregation agg = newAggregation(
    match(Criteria.where("status").is("PAID")),
    group("userId")
        .count().as("orderCount")
        .sum("amount").as("totalAmount")
        .max("createdAt").as("lastOrderAt"),
    sort(Sort.Direction.DESC, "totalAmount"),
    limit(20)
);

AggregationResults<UserOrderStat> results = mongoTemplate.aggregate(
    agg, "orders", UserOrderStat.class
);

五、索引设计 ​

5.1 单字段索引 ​

javascript
db.users.createIndex({ email: 1 })         // 1 升序,-1 降序(效果一样)
db.users.createIndex({ email: 1 }, { unique: true })  // 唯一索引

5.2 复合索引(顺序很重要) ​

javascript
db.orders.createIndex({ userId: 1, createdAt: -1 })

最左前缀:{ userId: 1, createdAt: -1 } 可用于:

  • { userId: 1 }(前缀)
  • { userId: 1, createdAt: -1 }(完整)
  • { userId: 1, createdAt: -1, status: 1 } 的 sort

但不能用于 { createdAt: -1 }(跳过了 userId)。

⚠️ 坑 4:乱序字段顺序会导致索引失效,慢查询排查时第一件事就是看执行计划。

5.3 覆盖索引 ​

查询的字段全部在索引里,ES / Mongo 直接从索引返回,不走文档。

javascript
db.users.createIndex({ name: 1, age: 1 })
db.users.find({ name: "tom" }, { name: 1, age: 1, _id: 0 })  // 覆盖,极快

5.4 看执行计划 ​

javascript
db.orders.find({ userId: 1001 }).explain("executionStats")

重点看 winningPlan.stage 和 totalDocsExamined vs nReturned:

  • totalDocsExamined 远大于 nReturned → 索引没起作用

六、副本集(Replica Set) ​

至少 1 主 2 从,提供高可用和读写分离。

bash
docker run -d --name mongo1 -p 27017:27017 mongo:7 --replSet rs0
docker run -d --name mongo2 -p 27018:27017 mongo:7 --replSet rs0
docker run -d --name mongo3 -p 27019:27017 mongo:7 --replSet rs0

# 进入 mongo1 初始化副本集
mongosh --port 27017
> rs.initiate({
    _id: "rs0",
    members: [
      { _id: 0, host: "localhost:27017" },
      { _id: 1, host: "localhost:27018" },
      { _id: 2, host: "localhost:27019" }
    ]
  })

⚠️ 坑 5:mongosh 默认读主节点,从节点读要加 ?readPreference=secondary 或 db.collection.find().readPref("secondary")。

七、分片集群(水平扩展) ​

数据量太大(>TB),单节点放不下 → 分片。

分片键选择:

选择影响
单调递增字段(_id)数据全在最新分片,热点
高基数随机字段(uid)数据均匀,但范围查询要 scatter-gather
复合分片键兼顾范围和均匀

八、本章小结 ​

要点关键
聚合管道$match → $group → $sort → $limit
$lookup等价 JOIN,但性能差,慎用
$unwind拆数组,每元素一条记录
复合索引最左前缀,字段顺序决定生死
覆盖索引索引完全包含查询字段,极快
副本集1 主 2 从,自动故障转移
分片键选择影响数据均匀和查询性能

动手练习 ​

  1. 写一个聚合管道,统计每个分类下的文章数和平均字数
  2. 给 users.email 加唯一索引,验证重复插入报错
  3. explain("executionStats") 看一条查询,确认是否走索引

下一章:第 12 章:RabbitMQ 进阶 →

本站基于 VitePress 构建 · 由 StackHub 团队维护