第 11 章:MongoDB 进阶
学习目标
- 掌握聚合管道($match、$group、$lookup)
- 设计合理的索引(单字段、复合、覆盖)
- 用副本集保证高可用
- 了解分片集群的水平扩展
一、聚合管道(Aggregation Pipeline)
聚合管道 = 一系列 stage,每个 stage 接收文档流,转换后传给下一个 stage。
javascript
db.orders.aggregate([
{ $match: { status: "PAID" } }, // 过滤
{ $group: { _id: "$userId", total: { $sum: "$amount" } } }, // 按用户分组求和
{ $sort: { total: -1 } }, // 排序
{ $limit: 10 } // 取前 10
])效果:每个付费用户的总消费额 Top 10。
二、常用 Stage
2.1 $match 过滤
javascript
{ $match: { status: "PAID", createdAt: { $gte: ISODate("2024-01-01") } } }⚠️ 坑 1:
$match尽量放在管道最前面,能利用索引,后面再 group。
2.2 $group 分组
javascript
{
$group: {
_id: "$category", // 分组键
count: { $sum: 1 }, // 计数
avgPrice: { $avg: "$price" }, // 平均
maxPrice: { $max: "$price" },
items: { $push: "$name" } // 收集到数组
}
}⚠️ 坑 2:
$push没限制会把所有文档字段塞进去,可能爆内存。生产用$push+$slice截断。
2.3 $project 投影
javascript
{
$project: {
name: 1, // 1 = 包含
price: 1,
isExpensive: { $gt: ["$price", 100] }
}
}2.4 $lookup JOIN
javascript
{
$lookup: {
from: "users",
localField: "userId",
foreignField: "_id",
as: "userInfo"
}
}效果:把 order 的 userId 关联到 users 的 _id,结果存在 userInfo 数组里(永远数组)。
⚠️ 坑 3:
$lookup性能比 MySQL JOIN 差很多,不要用在高频查询,先考虑反范式设计。
2.5 $unwind 拆数组
javascript
{ $unwind: "$items" } // 把 items: [a, b] 拆成两条:items: a 和 items: b三、聚合管道实战
需求:每个用户的订单数、总金额、最近订单时间。
javascript
db.orders.aggregate([
{ $match: { status: "PAID" } },
{
$group: {
_id: "$userId",
orderCount: { $sum: 1 },
totalAmount: { $sum: "$amount" },
lastOrderAt: { $max: "$createdAt" }
}
},
{ $sort: { totalAmount: -1 } },
{ $limit: 20 }
])四、Java 聚合
java
import static org.springframework.data.mongodb.core.aggregation.Aggregation.*;
Aggregation agg = newAggregation(
match(Criteria.where("status").is("PAID")),
group("userId")
.count().as("orderCount")
.sum("amount").as("totalAmount")
.max("createdAt").as("lastOrderAt"),
sort(Sort.Direction.DESC, "totalAmount"),
limit(20)
);
AggregationResults<UserOrderStat> results = mongoTemplate.aggregate(
agg, "orders", UserOrderStat.class
);五、索引设计
5.1 单字段索引
javascript
db.users.createIndex({ email: 1 }) // 1 升序,-1 降序(效果一样)
db.users.createIndex({ email: 1 }, { unique: true }) // 唯一索引5.2 复合索引(顺序很重要)
javascript
db.orders.createIndex({ userId: 1, createdAt: -1 })最左前缀:{ userId: 1, createdAt: -1 } 可用于:
{ userId: 1 }(前缀){ userId: 1, createdAt: -1 }(完整){ userId: 1, createdAt: -1, status: 1 }的 sort
但不能用于 { createdAt: -1 }(跳过了 userId)。
⚠️ 坑 4:乱序字段顺序会导致索引失效,慢查询排查时第一件事就是看执行计划。
5.3 覆盖索引
查询的字段全部在索引里,ES / Mongo 直接从索引返回,不走文档。
javascript
db.users.createIndex({ name: 1, age: 1 })
db.users.find({ name: "tom" }, { name: 1, age: 1, _id: 0 }) // 覆盖,极快5.4 看执行计划
javascript
db.orders.find({ userId: 1001 }).explain("executionStats")重点看 winningPlan.stage 和 totalDocsExamined vs nReturned:
totalDocsExamined远大于nReturned→ 索引没起作用
六、副本集(Replica Set)
至少 1 主 2 从,提供高可用和读写分离。
bash
docker run -d --name mongo1 -p 27017:27017 mongo:7 --replSet rs0
docker run -d --name mongo2 -p 27018:27017 mongo:7 --replSet rs0
docker run -d --name mongo3 -p 27019:27017 mongo:7 --replSet rs0
# 进入 mongo1 初始化副本集
mongosh --port 27017
> rs.initiate({
_id: "rs0",
members: [
{ _id: 0, host: "localhost:27017" },
{ _id: 1, host: "localhost:27018" },
{ _id: 2, host: "localhost:27019" }
]
})⚠️ 坑 5:
mongosh默认读主节点,从节点读要加?readPreference=secondary或db.collection.find().readPref("secondary")。
七、分片集群(水平扩展)
数据量太大(>TB),单节点放不下 → 分片。
分片键选择:
| 选择 | 影响 |
|---|---|
| 单调递增字段(_id) | 数据全在最新分片,热点 |
| 高基数随机字段(uid) | 数据均匀,但范围查询要 scatter-gather |
| 复合分片键 | 兼顾范围和均匀 |
八、本章小结
| 要点 | 关键 |
|---|---|
| 聚合管道 | $match → $group → $sort → $limit |
$lookup | 等价 JOIN,但性能差,慎用 |
$unwind | 拆数组,每元素一条记录 |
| 复合索引 | 最左前缀,字段顺序决定生死 |
| 覆盖索引 | 索引完全包含查询字段,极快 |
| 副本集 | 1 主 2 从,自动故障转移 |
| 分片键 | 选择影响数据均匀和查询性能 |
动手练习
- 写一个聚合管道,统计每个分类下的文章数和平均字数
- 给
users.email加唯一索引,验证重复插入报错 explain("executionStats")看一条查询,确认是否走索引
下一章:第 12 章:RabbitMQ 进阶 →