Skip to content
第 7 章 ⏱ 13 分钟阅读

第 7 章:ES 索引与映射 ​

学习目标 ​

  • 设计合理的 Mapping(字段类型、分词器)
  • 安装 IK 中文分词器
  • 理解 dynamic mapping 的坑
  • 掌握索引模板 / 别名

一、Mapping 是什么 ​

Mapping 就是 ES 的"表结构定义"——告诉 ES 每个字段什么类型、用什么分词器、能不能被搜索。和 MySQL 的 CREATE TABLE、TypeScript 的 interface 一个意思。

⚠️ Mapping 一旦设定,已存在的字段类型不能改(只能 reindex)。

精简例子:

bash
PUT /articles
{
  "mappings": {
    "properties": {
      "title":  { "type": "text" },      # 全文搜索,会分词
      "status": { "type": "keyword" }, # 精确匹配,不分词
      "views":  { "type": "integer" }, # 整数
      "created":{ "type": "date" } # 日期
    }
  }
}

二、字段类型速查 ​

类型用途示例
text全文搜索,会分词(按词搜)文章标题、内容
keyword精确匹配 / 排序 / 聚合(整体存)状态(0/1)、邮箱、标签
integer/long整数浏览数、价格
double/float浮点评分
date日期创建时间
boolean布尔是否删除
nested嵌套数组(数组里每项独立查)评论列表
object普通 JSON 对象作者 {name, age}

最常用的就两个:text(中文搜索)和 keyword(状态、标签)。其它类型按需选。

⚠️ 坑 1:数字字段不要用 text → 100 被当字符串,聚合 / 范围查询全失效。

⚠️ 坑 2:text 字段不能直接 sort / aggregate(需要开 fielddata,性能差);需要排序 / 聚合的字段用 keyword。

三、创建带 Mapping 的 Index ​

bash
PUT /articles
{
  "settings": {                                                // ① 索引系统配置
    "number_of_shards": 3,                                      //    主分片数 3
    "number_of_replicas": 1,                                    //    副本数 1
    "analysis": {                                               // ② 自定义分词器
      "analyzer": {
        "ik_smart_pinyin": {                                    //    名字(自定义)
          "type": "custom",                                     //    类型:自定义
          "tokenizer": "ik_smart"                              //    底层用 IK 智能分词
        }
      }
    }
  },
  "mappings": {                                                // ③ 字段类型定义
    "properties": {
      "id":        { "type": "keyword" },                       //    精确值(不分词)
      "title":     { "type": "text",                            //    全文搜索
                     "analyzer": "ik_max_word",                 //      写入:细切(词多)
                     "search_analyzer": "ik_smart" },            //      查询:粗切(词少)
      "content":   { "type": "text", "analyzer": "ik_max_word" }, //    内容:写入细切
      "status":    { "type": "keyword" },                       //    状态:精确值
      "views":     { "type": "integer" },                       //    浏览数:整数
      "price":     { "type": "scaled_float",                   //    价格:整数存小数(避浮点)
                     "scaling_factor": 100 },                   //      ×100 存储(9.99 存 999)
      "tags":      { "type": "keyword" },                       //    标签:精确值
      "createTime":{ "type": "date",                            //    时间
                     "format": "yyyy-MM-dd HH:mm:ss" },         //      自定义格式(默认只认 ISO)
      "author": {                                               //    嵌套对象
        "properties": {
          "name": { "type": "keyword" },                        //      作者名:精确
          "age":  { "type": "integer" }                         //      年龄:整数
        }
      }
    }
  }
}

⚠️ 坑 3:搜索时用 ik_smart(粗粒度,适合搜索),索引时用 ik_max_word(细粒度,提高召回)。

四、IK 中文分词器 ​

默认 standard 分词器对中文逐字切分,搜索体验极差。

安装 IK(8.x) ​

bash
# 进入 ES 容器
docker exec -it es bash
./bin/elasticsearch-plugin install https://github.com/medcl/elasticsearch-analysis-ik/releases/download/v8.11.0/elasticsearch-analysis-ik-8.11.0.zip
exit
docker restart es

验证 ​

bash
POST /_analyze
{ "analyzer": "ik_max_word", "text": "中华人民共和国国歌" }
# 输出:中华人民共和国 / 中华人民 / 中华 / 华人 / 人民共和国 / ...

POST /_analyze
{ "analyzer": "ik_smart", "text": "中华人民共和国国歌" }
# 输出:中华人民共和国 / 国歌

五、Dynamic Mapping ​

ES 默认开启 dynamic mapping,写入新字段会自动加进 mapping(但类型是猜的)。

bash
PUT /test/_doc/1 { "new_field": "hello" }
# 自动创建 new_field: text

三种策略 ​

策略行为
true(默认)新字段自动加入 mapping
false新字段存进 _source 但不可搜索
strict新字段直接报错

生产推荐 strict,避免脏数据污染 mapping。

怎么开 strict? ​

① 建索引时开 strict(整索引生效,推荐):

bash
PUT /articles
{
  "mappings": {
    "dynamic": "strict",                    # ← 严格模式:没声明的字段直接拒
    "properties": {
      "title": { "type": "text" },
      "views": { "type": "integer" }
    }
  }
}

效果验证:

bash
# 正常字段 OK
POST /articles/_doc/1
{ "title": "Java" }                         # ✅

# 多带未声明字段 → 直接报错
POST /articles/_doc/2
{ "title": "Python", "unknown_field": "x" } # ❌ strict_dynamic_mapping_exception

② 已有索引改为 strict(只能收紧):

bash
PUT /articles/_mapping
{
  "dynamic": "strict"
}
# 只能 true → false → strict,反过来不行

③ 单字段不开 strict,只关索引(存但搜不到):

bash
PUT /articles
{
  "mappings": {
    "properties": {
      "internal_note": {
        "type": "text",
        "index": false                      # ← 这个字段存进 _source,但搜不到
      }
    }
  }
}

⚠️ 坑 4:开发环境开了默认 true,某次上线传了 price: "abc" → 自动识别成 text,后续想改成 double 不行,只能 reindex。

六、修改 Mapping(实战) ​

6.1 加字段 ​

bash
PUT /articles/_mapping
{
  "properties": {
    "category": { "type": "keyword" }
  }
}

6.2 改字段类型(只能 reindex) ​

bash
# 1. 创建新 index 用新 mapping
PUT /articles_v2 { ... 新 mapping ... }                      # ① 建新索引(字段类型已存在不能改,只能新建)

# 2. reindex
POST /_reindex                                              # ② ES 自动搬运数据
{
  "source": { "index": "articles" },                         #    从旧索引
  "dest":   { "index": "articles_v2" }                      #    到新索引
}

# 3. 切换别名
POST /_aliases                                              # ③ 让业务访问从旧索引切到新索引
{
  "actions": [
    { "remove": { "index": "articles",    "alias": "articles_read" } },  # 摘下旧索引
    { "add":    { "index": "articles_v2", "alias": "articles_read" } }   # 挂上新索引
  ]
}

七、索引别名(零停机切换) ​

bash
POST /_aliases
{
  "actions": [
    { "add": { "index": "articles", "alias": "blog" } }
  ]
}

之后查询用 /blog/_search,底层 index 改了 alias 不变。

已经建好的索引,现在想加别名也来得及:

bash
POST /_aliases
{
  "actions": [
    { "add": { "index": "articles", "alias": "articles_read" } }
  ]
}

注意:加完别名还要把业务代码里的索引名改成别名(articles → articles_read),之后才能用别名做"无缝切换"。

⚠️ 坑 5:reindex 期间老 index 仍在写入,新 index 会缺数据 —— 要么停写,要么用别名双写一段时间。

八、索引模板(批量创建) ​

bash
PUT /_index_template/my_template
{
  "index_patterns": ["logs-*"],
  "template": {
    "settings": { "number_of_shards": 3 },
    "mappings": { "properties": { ... } }
  }
}

之后 PUT /logs-2024-01-01 自动套用模板。

九、本章小结 ​

要点关键
Mapping字段类型 + 分词器,写错就改不动
text vs keyword全文用 text,精确/聚合用 keyword
IK 分词中文搜索必备,ik_max_word 索引 / ik_smart 搜索
Dynamic生产推荐 strict,避免脏数据
改字段加字段直接改,改类型必须 reindex
别名零停机切换,reindex + alias 双写

动手练习 ​

  1. 创建 news index,字段包含 title(text+ik)、status(keyword)、publishDate(date)
  2. 用 _analyze 验证 IK 分词效果
  3. 试一次"加字段 → 查 mapping" → "改字段类型失败 → reindex" 的完整流程

下一章:第 8 章:ES 查询 DSL →

本站基于 VitePress 构建 · 由 StackHub 团队维护