第 7 章:ES 索引与映射
学习目标
- 设计合理的 Mapping(字段类型、分词器)
- 安装 IK 中文分词器
- 理解 dynamic mapping 的坑
- 掌握索引模板 / 别名
一、Mapping 是什么
Mapping 就是 ES 的"表结构定义"——告诉 ES 每个字段什么类型、用什么分词器、能不能被搜索。和 MySQL 的 CREATE TABLE、TypeScript 的 interface 一个意思。
⚠️ Mapping 一旦设定,已存在的字段类型不能改(只能 reindex)。
精简例子:
PUT /articles
{
"mappings": {
"properties": {
"title": { "type": "text" }, # 全文搜索,会分词
"status": { "type": "keyword" }, # 精确匹配,不分词
"views": { "type": "integer" }, # 整数
"created":{ "type": "date" } # 日期
}
}
}二、字段类型速查
| 类型 | 用途 | 示例 |
|---|---|---|
text | 全文搜索,会分词(按词搜) | 文章标题、内容 |
keyword | 精确匹配 / 排序 / 聚合(整体存) | 状态(0/1)、邮箱、标签 |
integer/long | 整数 | 浏览数、价格 |
double/float | 浮点 | 评分 |
date | 日期 | 创建时间 |
boolean | 布尔 | 是否删除 |
nested | 嵌套数组(数组里每项独立查) | 评论列表 |
object | 普通 JSON 对象 | 作者 {name, age} |
最常用的就两个:
text(中文搜索)和keyword(状态、标签)。其它类型按需选。⚠️ 坑 1:数字字段不要用
text→100被当字符串,聚合 / 范围查询全失效。⚠️ 坑 2:
text字段不能直接 sort / aggregate(需要开fielddata,性能差);需要排序 / 聚合的字段用keyword。
三、创建带 Mapping 的 Index
PUT /articles
{
"settings": { // ① 索引系统配置
"number_of_shards": 3, // 主分片数 3
"number_of_replicas": 1, // 副本数 1
"analysis": { // ② 自定义分词器
"analyzer": {
"ik_smart_pinyin": { // 名字(自定义)
"type": "custom", // 类型:自定义
"tokenizer": "ik_smart" // 底层用 IK 智能分词
}
}
}
},
"mappings": { // ③ 字段类型定义
"properties": {
"id": { "type": "keyword" }, // 精确值(不分词)
"title": { "type": "text", // 全文搜索
"analyzer": "ik_max_word", // 写入:细切(词多)
"search_analyzer": "ik_smart" }, // 查询:粗切(词少)
"content": { "type": "text", "analyzer": "ik_max_word" }, // 内容:写入细切
"status": { "type": "keyword" }, // 状态:精确值
"views": { "type": "integer" }, // 浏览数:整数
"price": { "type": "scaled_float", // 价格:整数存小数(避浮点)
"scaling_factor": 100 }, // ×100 存储(9.99 存 999)
"tags": { "type": "keyword" }, // 标签:精确值
"createTime":{ "type": "date", // 时间
"format": "yyyy-MM-dd HH:mm:ss" }, // 自定义格式(默认只认 ISO)
"author": { // 嵌套对象
"properties": {
"name": { "type": "keyword" }, // 作者名:精确
"age": { "type": "integer" } // 年龄:整数
}
}
}
}
}⚠️ 坑 3:搜索时用
ik_smart(粗粒度,适合搜索),索引时用ik_max_word(细粒度,提高召回)。
四、IK 中文分词器
默认 standard 分词器对中文逐字切分,搜索体验极差。
安装 IK(8.x)
# 进入 ES 容器
docker exec -it es bash
./bin/elasticsearch-plugin install https://github.com/medcl/elasticsearch-analysis-ik/releases/download/v8.11.0/elasticsearch-analysis-ik-8.11.0.zip
exit
docker restart es验证
POST /_analyze
{ "analyzer": "ik_max_word", "text": "中华人民共和国国歌" }
# 输出:中华人民共和国 / 中华人民 / 中华 / 华人 / 人民共和国 / ...
POST /_analyze
{ "analyzer": "ik_smart", "text": "中华人民共和国国歌" }
# 输出:中华人民共和国 / 国歌五、Dynamic Mapping
ES 默认开启 dynamic mapping,写入新字段会自动加进 mapping(但类型是猜的)。
PUT /test/_doc/1 { "new_field": "hello" }
# 自动创建 new_field: text三种策略
| 策略 | 行为 |
|---|---|
true(默认) | 新字段自动加入 mapping |
false | 新字段存进 _source 但不可搜索 |
strict | 新字段直接报错 |
生产推荐 strict,避免脏数据污染 mapping。
怎么开 strict?
① 建索引时开 strict(整索引生效,推荐):
PUT /articles
{
"mappings": {
"dynamic": "strict", # ← 严格模式:没声明的字段直接拒
"properties": {
"title": { "type": "text" },
"views": { "type": "integer" }
}
}
}效果验证:
# 正常字段 OK
POST /articles/_doc/1
{ "title": "Java" } # ✅
# 多带未声明字段 → 直接报错
POST /articles/_doc/2
{ "title": "Python", "unknown_field": "x" } # ❌ strict_dynamic_mapping_exception② 已有索引改为 strict(只能收紧):
PUT /articles/_mapping
{
"dynamic": "strict"
}
# 只能 true → false → strict,反过来不行③ 单字段不开 strict,只关索引(存但搜不到):
PUT /articles
{
"mappings": {
"properties": {
"internal_note": {
"type": "text",
"index": false # ← 这个字段存进 _source,但搜不到
}
}
}
}⚠️ 坑 4:开发环境开了默认
true,某次上线传了price: "abc"→ 自动识别成text,后续想改成double不行,只能 reindex。
六、修改 Mapping(实战)
6.1 加字段
PUT /articles/_mapping
{
"properties": {
"category": { "type": "keyword" }
}
}6.2 改字段类型(只能 reindex)
# 1. 创建新 index 用新 mapping
PUT /articles_v2 { ... 新 mapping ... } # ① 建新索引(字段类型已存在不能改,只能新建)
# 2. reindex
POST /_reindex # ② ES 自动搬运数据
{
"source": { "index": "articles" }, # 从旧索引
"dest": { "index": "articles_v2" } # 到新索引
}
# 3. 切换别名
POST /_aliases # ③ 让业务访问从旧索引切到新索引
{
"actions": [
{ "remove": { "index": "articles", "alias": "articles_read" } }, # 摘下旧索引
{ "add": { "index": "articles_v2", "alias": "articles_read" } } # 挂上新索引
]
}七、索引别名(零停机切换)
POST /_aliases
{
"actions": [
{ "add": { "index": "articles", "alias": "blog" } }
]
}之后查询用 /blog/_search,底层 index 改了 alias 不变。
已经建好的索引,现在想加别名也来得及:
POST /_aliases
{
"actions": [
{ "add": { "index": "articles", "alias": "articles_read" } }
]
}注意:加完别名还要把业务代码里的索引名改成别名(articles → articles_read),之后才能用别名做"无缝切换"。
⚠️ 坑 5:reindex 期间老 index 仍在写入,新 index 会缺数据 —— 要么停写,要么用别名双写一段时间。
八、索引模板(批量创建)
PUT /_index_template/my_template
{
"index_patterns": ["logs-*"],
"template": {
"settings": { "number_of_shards": 3 },
"mappings": { "properties": { ... } }
}
}之后 PUT /logs-2024-01-01 自动套用模板。
九、本章小结
| 要点 | 关键 |
|---|---|
| Mapping | 字段类型 + 分词器,写错就改不动 |
| text vs keyword | 全文用 text,精确/聚合用 keyword |
| IK 分词 | 中文搜索必备,ik_max_word 索引 / ik_smart 搜索 |
| Dynamic | 生产推荐 strict,避免脏数据 |
| 改字段 | 加字段直接改,改类型必须 reindex |
| 别名 | 零停机切换,reindex + alias 双写 |
动手练习
- 创建
newsindex,字段包含title(text+ik)、status(keyword)、publishDate(date) - 用
_analyze验证 IK 分词效果 - 试一次"加字段 → 查 mapping" → "改字段类型失败 → reindex" 的完整流程
下一章:第 8 章:ES 查询 DSL →