Skip to content

Wildcard query normalizer mishandles escapes and re-escaping, breaking literal search for any character that normalizes to or from * ? \ #150699

Description

@TheRiffRafi

Elasticsearch Version

Reproduced on main. The relevant code is unchanged across recent 8.x and 9.x releases.

Installed Plugins

analysis-icu (for the real-world trigger). The reproduction below uses a built-in pattern_replace char filter and needs no plugins.

Java Version

N/A (logic bug, version independent)

OS Version

N/A

Problem Description

When a keyword field uses a normalizer whose output adds, removes, or rewrites the wildcard control characters *, ?, or \, wildcard queries return incorrect results. The most common trigger is the ICU Analysis plugin's NFKC normalization, which maps fullwidth East Asian forms to ASCII (* to *, ? to ?, \ to \). The same defect applies to any normalizer that maps in either direction across these characters.

There are two distinct, direction-dependent failure modes:

  • Bug 1 (escape contents bypass normalization). A literal search written with an escape, for example \*, returns zero hits. The whole \X unit is appended verbatim, so the escaped character is never normalized. The query then looks for a character that does not exist in the index.
  • Bug 2 (normalizer output is not re-escaped). A literal search written without an escape, for example a bare fullwidth *, is normalized to * and then treated as a wildcard operator. A character that was data silently becomes an operator and matches far more documents than intended. With ? it becomes a single-character wildcard; with fullwidth \ the resulting stray \ typically produces zero matches.
    Which bug you hit depends on the normalizer direction. A normalizer that maps * to * exhibits Bug 1. A normalizer that maps * to * (ICU NFKC) exhibits Bug 2. As a result there is no normalizer that rewrites to or from these base forms for which literal wildcard search behaves correctly, and no query formulation that recovers correct behavior (see the impact table below).

Root cause

The normalization-aware wildcard path is StringFieldType.normalizeWildcardPattern(...), called from both KeywordFieldMapper wildcard query paths. It segments the query with:

// server/src/main/java/org/elasticsearch/index/mapper/StringFieldType.java
private static final Pattern WILDCARD_PATTERN = Pattern.compile("(\\.)|([?*]+)");

Group 1 matches an escape sequence \X. Group 2 matches runs of literal ?/*. The loop normalizes the gaps between matches and appends each match verbatim:

while (wildcardMatcher.find()) {
    if (wildcardMatcher.start() > 0) {
        String chunk = value.substring(last, wildcardMatcher.start());
        BytesRef normalized = normalizer.normalize(fieldname, chunk);
        sb.append(normalized);                 // (A) gaps normalized, output not re-escaped
    }
    sb.append(new BytesRef(wildcardMatcher.group()));  // (B) match appended verbatim, contents not normalized
    last = wildcardMatcher.end();
}
if (last < value.length()) {
    BytesRef normalized = normalizer.normalize(fieldname, value.substring(last));
    sb.append(normalized);                     // (A) trailing gap, same gap handling
}

The in-code comment states the intent is to "normalize everything except wildcard characters." That intent is implemented only for the operator-preservation case. The two failure modes map directly onto the marked lines:

  • Bug 1 is line (B). The escape match \X is appended as-is, so the escaped character is never passed to normalizer.normalize(...).
  • Bug 2 is line (A). A character that is not an ASCII ?/* falls into a normalized gap. The normalizer may emit *, ?, or \, and that output is appended without re-escaping, so emitted operators are interpreted as wildcards.

Steps to Reproduce

The pattern_replace char filter below mimics ICU NFKC for *, ?, and \ in isolation. Behavior is identical with the ICU plugin.

PUT /icu-wildcard-bug-test
{
  "settings": {
    "analysis": {
      "normalizer": {
        "icu_nfkc_mimic": {
          "type": "custom",
          "char_filter": ["fw_star", "fw_question", "fw_backslash"]
        }
      },
      "char_filter": {
        "fw_star":     { "type": "pattern_replace", "pattern": "*", "replacement": "*" },
        "fw_question": { "type": "pattern_replace", "pattern": "?", "replacement": "?" },
        "fw_backslash":{ "type": "pattern_replace", "pattern": "\", "replacement": "\\\\" }
      }
    }
  },
  "mappings": {
    "properties": {
      "label":         { "type": "keyword" },
      "kw_vanilla":    { "type": "keyword" },
      "kw_normalized": { "type": "keyword", "normalizer": "icu_nfkc_mimic" }
    }
  }
}
POST /icu-wildcard-bug-test/_bulk
{ "index": { "_id": "1" } }
{ "label": "fw-star",     "kw_vanilla": "東京*大阪", "kw_normalized": "東京*大阪" }
{ "index": { "_id": "2" } }
{ "label": "fw-question", "kw_vanilla": "東京?大阪", "kw_normalized": "東京?大阪" }
{ "index": { "_id": "3" } }
{ "label": "fw-backslash","kw_vanilla": "東京\大阪", "kw_normalized": "東京\大阪" }
{ "index": { "_id": "4" } }
{ "label": "no-special",  "kw_vanilla": "東京大阪",   "kw_normalized": "東京大阪" }

Index-time normalization is correct. Confirmed stored values: 東京*大阪 to 東京*大阪, 東京?大阪 to 東京?大阪, 東京\大阪 to 東京\大阪, 東京大阪 to 東京大阪.

Bug 1, escape contents bypassed. A user typing the fullwidth form and escaping it for a literal match gets zero hits:

GET /icu-wildcard-bug-test/_search
{ "query": { "wildcard": { "kw_normalized": { "value": "東京\\*大阪" } } } }
Query value Field Expected Actual
東京\*大阪 kw_normalized 1 (fw-star) 0
東京\*大阪 kw_normalized 1 (fw-star) 1 (only works if caller types the post-normalization ASCII form)
東京\*大阪 kw_vanilla 1 (fw-star) 1

Bug 2, output not re-escaped. A user searching the fullwidth literal without an escape gets wildcard behavior:

GET /icu-wildcard-bug-test/_search
{ "query": { "wildcard": { "kw_normalized": { "value": "東京*大阪" } } } }
Query value Expected Actual Matched
東京*大阪 1 (fw-star, literal *) 4 all docs
東京?大阪 1 (fw-question, literal ?) 3 fw-star, fw-question, fw-backslash
東京*大阪 4 (operator *) 4 all docs (confirms operator path is correct)

Impact: no valid query formulation exists

User intent Query Result Cause
Literal * via fullwidth form 東京*大阪 wrong, wildcard behavior Bug 2
Literal * via fullwidth escape 東京\*大阪 wrong, 0 results Bug 1
Literal * via ASCII escape 東京\*大阪 correct, but defeats normalization caller must know the post-normalization form
* as operator 東京* correct operator path is fine

The only working formulation forces the caller to be normalizer aware and type the post-normalization ASCII form, which defeats the purpose of normalization on search terms. For fullwidth \ the failure is worse, since the stray \ generally yields zero matches rather than an over-inclusive set, so affected users silently get no results.

Expected behavior

The parser already segments the query into typed tokens (escapes and operator runs), so the infrastructure is in place. Two handling changes are needed inside normalizeWildcardPattern:

  1. For an escape match \X: strip the \, normalize X as a literal, re-escape if the normalized output is *, ?, or \, then append. The \ is an instruction to treat the next character as literal, not data to preserve unmodified.
  2. For normalized gaps: after normalizing a chunk, escape any *, ?, or \ that the normalizer produced, so emitted operators are treated as the literal data they came from.
    Operator runs in group 2 should continue to be preserved as-is, which is the existing correct behavior.

Related issues

Logs (if relevant)

No response

Activity

  1. added
    :Search Foundations/MappingIndex mappings, including merging and defining field types
    Team:Search FoundationsMeta label for the Search Foundations team in Elasticsearch
    and removed
    needs:triageRequires assignment of a team area label
    on Jun 4, 2026
  2. elasticsearchmachine commented on Jun 4, 2026

    @elasticsearchmachine
    Collaborator

    Pinging @elastic/es-search-foundations (Team:Search Foundations)

  3. fcofdez commented on Jun 4, 2026

    @fcofdez
    Contributor

    I'm not 100% sure if Search foundations is the right team to handle this, but it looks like this falls into mappings?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    :Search Foundations/MappingIndex mappings, including merging and defining field types>bugTeam:Search FoundationsMeta label for the Search Foundations team in Elasticsearch

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions