带你读《Elastic Stack 实战手册》之34:——3.4.2.17.3.全文搜索/精确搜索(9)

简介: 带你读《Elastic Stack 实战手册》之34:——3.4.2.17.3.全文搜索/精确搜索(9)


《Elastic Stack 实战手册》——三、产品能力——3.4.入门篇——3.4.2.Elasticsearch基础应用——3.4.2.17.Text analysis, settings 及 mappings——3.4.2.17.3.全文搜索/精确搜索(8) https://developer.aliyun.com/article/1229934



四、基于全文的查询方法

 

基于全文的方法主要有:match/match_phrase/match_phrase_prefix/multi_match/match_bool_prefix/query_string/simple_query_string/intervals/combined_fields 九种方法。

 

在全文查询的复杂方法中,很多基于 match 查询的参数,如:analyzer, boost, operator, minimum_should_match, fuzziness, lenient, prefix_length, max_expansions, fuzzy_rewrite,zero_terms_query 和 cutoff_frequency 都能在其它方法中使用。

 

4.1 match

 

match 查询是一个基础的全文搜索方法,它会将查询的短语进行分词后对某一字段进行查询。match 查询对被分词后的 token 并没有强顺序关系,只要匹配就可以返回。


使用方法:

PUT my-index-000001
POST my-index-000001/_mapping
{"properties":{"message":{"type":"text"}}}
POST my-index-000001/_bulk
{ "index": { "_id": 1 }}
{ "message": "this is a test" }
GET my-index-000001/_search
{
  "query": {
    "match": {
      "message": {
        "query": "this is a test"
      }
    }
  }
}

参数:

 

l analyzer:设置将查询的短语转换成 token 的分词器。默认与字段索引时的分词器一致,即 mapping 设置的 analyzer 参数。如果没有设置,则使用默认的分词。

l auto_generate_synonyms_phrase_query:如果为 true,为多项同义词自动生成短语查询。这个与分词器中设置的同义词相关。默认值为 true。

l fuzziness:允许匹配的最大编辑距离。可以是 0/1/2/AUTO

l max_expansions:创建的最大变体或者扩展词项数。默认为 50。

l prefix_length:在创建展开时保持不变的起始字符数。默认值为 0。

l fuzzy_transpositions:指编辑是否包括两个相邻字符的换位( ab → ba )。默认值为 true。

l lenient:如果为真,则忽略基于格式的错误,例如为数字字段提供 text 查询值。默认值为 false。

l operator:查询文本中的布尔逻辑。

○ OR:默认值。例如,将查询值 “capital of Hungary” 解释为 “capital” 或者 “of” 或者 “Hungary”。

○ AND:例如,将查询值 “capital of Hungary” 解释为 “capital” 和 “of” 和 “Hungary”。

l minimum_should_match:要返回的文档必须匹配的最小词项数。例如,“capital of Hungary” 被分词成 “capital”、“of”、“Hungary” 三个词项,minimum_should_match 设置为 2,则文档必须匹配前面三个词项中的两个才能返回。

l zero_terms_query:指示如果分析器删除所有词项时(例如使用停顿词分词器时),是否不返回文档。

○ none:默认值。如果分词器删除所有词项时,则不返回文档。

○ all:与none相反,返回所有文档。

 

相关使用方法:

 

match 查询的 operator 和 minimum_should_match

 

先创建一个测试索引和相关测试数据:


PUT my-index-000001
POST my-index-000001/_mapping
{
  "properties": {
    "message": {
      "type": "text"
    }
  }
}
PUT my-index-000001/_doc/1
{ "message":"this is test"}
PUT my-index-000001/_doc/2
{ "message":"this is a test again"}
PUT my-index-000001/_doc/3
{ "message":"this is  not a test"}

使用默认的分词器,可以看到几个文档会被解析成 "this"、"is"、"a"、"test"、"again"、"not" 这几个词项。

 

POST _analyze
{
  "text": [
    "this is test",
    "this is a test again",
    "this is  not a test"
  ]
}
# 返回结果
{
  "tokens" : [
    {
      "token" : "this",
      "start_offset" : 0,
      "end_offset" : 4,
      "type" : "<ALPHANUM>",
      "position" : 0
    },
    {
      "token" : "is",
      "start_offset" : 5,
      "end_offset" : 7,
      "type" : "<ALPHANUM>",
      "position" : 1
},
   {
      "token" : "test",
      "start_offset" : 8,
      "end_offset" : 12,
      "type" : "<ALPHANUM>",
      "position" : 2
    },
    {
      "token" : "this",
      "start_offset" : 13,
      "end_offset" : 17,
      "type" : "<ALPHANUM>",
      "position" : 3
    },
    {
      "token" : "is",
      "start_offset" : 18,
      "end_offset" : 20,
      "type" : "<ALPHANUM>",
      "position" : 4
    },
    {
      "token" : "a",
      "start_offset" : 21,
      "end_offset" : 22,
      "type" : "<ALPHANUM>",
      "position" : 5
    },
    {
      "token" : "test",
      "start_offset" : 23,
      "end_offset" : 27,
      "type" : "<ALPHANUM>",
        "position" : 6
    },
    {
      "token" : "again",
      "start_offset" : 28,
      "end_offset" : 33,
      "type" : "<ALPHANUM>",
      "position" : 7
    },
    {
      "token" : "this",
      "start_offset" : 34,
      "end_offset" : 38,
      "type" : "<ALPHANUM>",
      "position" : 8
    },
    {
      "token" : "is",
      "start_offset" : 39,
      "end_offset" : 41,
      "type" : "<ALPHANUM>",
      "position" : 9
    },
    {
      "token" : "not",
      "start_offset" : 43,
      "end_offset" : 46,
      "type" : "<ALPHANUM>",
      "position" : 10
    },
    {
      "token" : "a",
      "start_offset" : 47,
      "end_offset" : 48,
      "type" : "<ALPHANUM>",
      "position" : 11
    },
    {
      "token" : "test",
      "start_offset" : 49,
      "end_offset" : 53,
      "type" : "<ALPHANUM>",
      "position" : 12
    }
  ]
}

match 查询 "this is a test",这个短语也会被分词成 "this"、"is"、"a"、"test"


POST _analyze
{
  "text": [
    "this is a test"
  ]
}
# 返回结果
{
  "tokens" : [
    {
      "token" : "this",
      "start_offset" : 0,
      "end_offset" : 4,
      "type" : "<ALPHANUM>",
      "position" : 0
    },
{
      "token" : "is",
      "start_offset" : 5,
      "end_offset" : 7,
      "type" : "<ALPHANUM>",
      "position" : 1
    },
    {
      "token" : "a",
      "start_offset" : 8,
      "end_offset" : 9,
      "type" : "<ALPHANUM>",
      "position" : 2
    },
    {
      "token" : "test",
      "start_offset" : 10,
      "end_offset" : 14,
      "type" : "<ALPHANUM>",
      "position" : 3
    }
  ]
}

如果按照 match 默认 or 的查询逻辑,只要有一个词项匹配就会返回,那么测试的三个文档将全部返回。


GET /_search
{
  "query": {
    "match": {
      "message": {
        "query": "this is a test",
        "operator": "or"
      }
    }
  }
}
# 返回结果
{
  "took" : 16,
  "timed_out" : false,
  "_shards" : {
    "total" : 48,
    "successful" : 48,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 3,
      "relation" : "eq"
    },
    "max_score" : 0.5298672,
    "hits" : [
      {
        "_index" : "my-index-000001",
        "_type" : "_doc",
        "_id" : "3",
        "_score" : 0.5298672,
        "_source" : {
          "message" : "this is  not a test"
        }
      },
      {
        "_index" : "my-index-000001",
        "_type" : "_doc",
        "_id" : "2",
        "_score" : 0.5298672,
        "_source" : {
          "message" : "this is a test again"
        }
      },
      {
        "_index" : "my-index-000001",
        "_type" : "_doc",
        "_id" : "1",
        "_score" : 0.30433932,
        "_source" : {
          "message" : "this is test"
        }
      }
    ]
  }
}


但是如果设置 minimum_should_match 为 4,则需要有四个词项匹配,那么只有文档 2 和 3 符合了。


GET /_search
{
  "query": {
    "match": {
      "message": {
        "query": "this is a test",
        "operator": "or",
        "minimum_should_match": 4
      }
}
  }
}
# 返回结果
{
  ......
   "hits" : {
    "total" : {
      "value" : 2,
      "relation" : "eq"
    },
    "max_score" : 0.5298672,
    "hits" : [
      {
        "_index" : "my-index-000001",
        "_type" : "_doc",
        "_id" : "3",
        "_score" : 0.5298672,
        "_source" : {
          "message" : "this is  not a test"
        }
      },
      {
        "_index" : "my-index-000001",
        "_type" : "_doc",
        "_id" : "2",
        "_score" : 0.5298672,
        "_source" : {
          "message" : "this is a test again"
        }
      }
    ]
  }
}


然后,再看一下 operator 为 and 的时候。其实,可以发现 and 的情况与之前设置 minimum_should_match 为 4 的查询一致,因为两者都代表查询时,每个词项都需要匹配上。


GET /_search
{
  "query": {
    "match": {
      "message": {
        "query": "this is a test",
        "operator": "and"
      }
    }
  }
}
# 返回结果
{
 ......
  "hits" : {
    "total" : {
      "value" : 2,
      "relation" : "eq"
    },
    "max_score" : 0.5298672,
    "hits" : [
      {
        "_index" : "my-index-000001",
        "_type" : "_doc",
        "_id" : "3",
        "_score" : 0.5298672,
        "_source" : {
          "message" : "this is  not a test"
        }
      },
      {
        "_index" : "my-index-000001",
        "_type" : "_doc",
        "_id" : "2",
        "_score" : 0.5298672,
        "_source" : {
          "message" : "this is a test again"
        }
      }
    ]
  }
}

match 中的模糊查询

 

fuzziness 的一系列参数可以使 match 解析出的词项进行模糊匹配。具体相关参数的使用方法与 fuzzy 查询一致,因此不详细展开了。

 

来看下面的例子:


GET /_search
{
  "query": {
    "match": {
      "message": {
        "query": "this is a test",
        "fuzziness": "auto"
        , "operator": "and"
      }
    }
  }
}
# 返回结果
{
  ......
  "hits" : {
    "total" : {
      "value" : 2,
      "relation" : "eq"
    },
    "max_score" : 0.50886154,
    "hits" : [
      {
        "_index" : "my-index-000001",
        "_type" : "_doc",
        "_id" : "3",
        "_score" : 0.50886154,
        "_source" : {
          "message" : "this is  not a test"
        }
      },
      {
        "_index" : "my-index-000001",
        "_type" : "_doc",
        "_id" : "2",
        "_score" : 0.50886154,
        "_source" : {
          "message" : "this is a test again"
        }
      }
    ]
  }
}


很明显,虽然查询的内容中错误的把 “test” 写成了 “testt”,但是经过 fuzziness 参数的调整,达到了纠错的效果。



《Elastic Stack 实战手册》——三、产品能力——3.4.入门篇——3.4.2.Elasticsearch基础应用——3.4.2.17.Text analysis, settings 及 mappings——3.4.2.17.3.全文搜索/精确搜索(10) https://developer.aliyun.com/article/1229931

 

相关实践学习
以电商场景为例搭建AI语义搜索应用
本实验旨在通过阿里云Elasticsearch结合阿里云搜索开发工作台AI模型服务,构建一个高效、精准的语义搜索系统,模拟电商场景,深入理解AI搜索技术原理并掌握其实现过程。
ElasticSearch 最新快速入门教程
本课程由千锋教育提供。全文搜索的需求非常大。而开源的解决办法Elasricsearch(Elastic)就是一个非常好的工具。目前是全文搜索引擎的首选。本系列教程由浅入深讲解了在CentOS7系统下如何搭建ElasticSearch,如何使用Kibana实现各种方式的搜索并详细分析了搜索的原理,最后讲解了在Java应用中如何集成ElasticSearch并实现搜索。 &nbsp;
相关文章
|
存储 编解码 数据可视化
低代码多分支协同开发的建设与实践
低代码多分支协同开发的建设与实践
1795 0
低代码多分支协同开发的建设与实践
|
存储 Kubernetes Linux
带你读《存储漫谈Ceph原理与实践》第三章接入层3.1块存储 RBD
《存储漫谈Ceph原理与实践》第三章接入层3.1块存储 RBD
带你读《存储漫谈Ceph原理与实践》第三章接入层3.1块存储 RBD
|
8月前
|
弹性计算 安全 关系型数据库
阿里云产品组合购买套餐解析:特惠方案、组合配置与套餐价格参考
2026年阿里云继续推出多样化云产品组合特惠方案,以99元、199元云服务器为核心,适配个人开发者、中小企业等多类用户需求。99元经济型e实例(2核2G+3M带宽+40G云盘)适合简单网站、开发测试;199元u1实例(2核4G+5M带宽+80G云盘)支持企业级应用。方案包含域名+ECS+智能建站、ECS+RDS、ECS+云安全中心等组合,提供自动扩缩容、跨可用区容灾等技术保障,结合续费同价、组合优惠等政策,助力用户实现从基础建站到企业级应用的一站式上云。
747 3
|
11月前
|
存储 容灾 安全
阿里云怎么配置跨区域复制?
在数据为王时代,阿里云OSS跨区域复制助力企业实现异地备份与容灾。本文手把手教你配置增量数据自动同步,保障业务连续性与数据安全。涵盖应用场景、核心优势及详细操作步骤,助你轻松上云。
|
5月前
|
存储 运维 安全
阿里云盘企业版支持哪些功能?限速吗?支持人数及存储费用清单(附2026最新价格表)
阿里云盘企业版CDE 2026年重磅升级,官网:https://t.aliyun.com/U/wbANEa 人数5人起步,500GB仅169.9元/年,不限速、不限流量;支持在线编辑、智能检索、水印预览、SSO单点登录及金融级安全。含存储、外流、AI处理全功能,无隐性费用。
|
6月前
|
传感器 人工智能 监控
协作机器人和工业机器人的区别
协作机器人(Cobot)是专为人机协同设计的工业机器人分支,以安全、灵活、易用为核心,通过力控感知、速度监控与ISO/TS 15066认证实现无围栏共作;支持拖拽示教、快速换型,部署快、成本低、ROI短(6–18个月),适用于打磨、柔性装配、医疗辅助等非标场景。(239字)
|
6月前
|
缓存 监控 算法
详细介绍一下淘宝商品详情API接口系列的调用频率限制
淘宝商品详情 API 系列(核心为taobao.item.get)的调用频率限制,采用应用级 QPS + 分钟级 + 日总量三重管控,按开发者身份、应用类型、接口版本、套餐等级分级执行,是平台保障服务稳定性、防滥用的核心机制。以下从限制规则、限流机制、超限处理、提额申请、调用方控频方案五方面详细说明。
|
6月前
|
人工智能 前端开发 Serverless
OpenClaw 实战:让AI 页面“秒开即用”,实现 Vibecoding 真正闭环
函数计算AgentRun Sandbox推出多端口沙箱能力,秒级冷启动、免运维、支持双模式路由,完美解决AI生成代码“写完即跑、所见即所得”的预览难题。结合OpenClaw,可实现对话中自动生成HTML/React页面、云端运行并返回实时预览链接,打通Vibecoding最后一公里。
|
6月前
|
供应链 数据处理 数据库
医疗器械唯一标识(UDI)赋码产线升级改造全方案
本方案面向医疗器械企业,聚焦法规合规、产线适配、自动赋码、数据互通与质量可控五大核心,提供覆盖小/中/大批量产能的标准化UDI产线改造方案,支持人工、半自动、全自动三种模式,兼容多包装类型,满足YY/T 1879、1943等标准及国家UDI实施要求。(239字)
627 0
|
8月前
|
人工智能 弹性计算 运维
阿里云无影云电脑商业版是什么?优势、应用场景及业务模式全解析
阿里云无影云电脑商业版是面向直播、AI教育、设计渲染等高价值场景的云上工作空间,融合“硬件+软件+安全+运维”能力,提供开箱即用的专业环境。具备极致体验、原生安全、统一运维、绿色节能与AI智能五大优势,支持六大行业场景应用,助力企业降本增效、合规创新,重塑数字化生产力。

热门文章

最新文章