Elasticsearch Version
Reproduced on main. The relevant code is unchanged across recent 8.x and 9.x releases.
Installed Plugins
analysis-icu (for the real-world trigger). The reproduction below uses a built-in pattern_replace char filter and needs no plugins.
Java Version
N/A (logic bug, version independent)
OS Version
N/A
Problem Description
When a keyword field uses a normalizer whose output adds, removes, or rewrites the wildcard control characters *, ?, or \, wildcard queries return incorrect results. The most common trigger is the ICU Analysis plugin's NFKC normalization, which maps fullwidth East Asian forms to ASCII (* to *, ? to ?, \ to \). The same defect applies to any normalizer that maps in either direction across these characters.
There are two distinct, direction-dependent failure modes:
- Bug 1 (escape contents bypass normalization). A literal search written with an escape, for example
\*, returns zero hits. The whole \X unit is appended verbatim, so the escaped character is never normalized. The query then looks for a character that does not exist in the index.
- Bug 2 (normalizer output is not re-escaped). A literal search written without an escape, for example a bare fullwidth
*, is normalized to * and then treated as a wildcard operator. A character that was data silently becomes an operator and matches far more documents than intended. With ? it becomes a single-character wildcard; with fullwidth \ the resulting stray \ typically produces zero matches.
Which bug you hit depends on the normalizer direction. A normalizer that maps * to * exhibits Bug 1. A normalizer that maps * to * (ICU NFKC) exhibits Bug 2. As a result there is no normalizer that rewrites to or from these base forms for which literal wildcard search behaves correctly, and no query formulation that recovers correct behavior (see the impact table below).
Root cause
The normalization-aware wildcard path is StringFieldType.normalizeWildcardPattern(...), called from both KeywordFieldMapper wildcard query paths. It segments the query with:
// server/src/main/java/org/elasticsearch/index/mapper/StringFieldType.java
private static final Pattern WILDCARD_PATTERN = Pattern.compile("(\\.)|([?*]+)");
Group 1 matches an escape sequence \X. Group 2 matches runs of literal ?/*. The loop normalizes the gaps between matches and appends each match verbatim:
while (wildcardMatcher.find()) {
if (wildcardMatcher.start() > 0) {
String chunk = value.substring(last, wildcardMatcher.start());
BytesRef normalized = normalizer.normalize(fieldname, chunk);
sb.append(normalized); // (A) gaps normalized, output not re-escaped
}
sb.append(new BytesRef(wildcardMatcher.group())); // (B) match appended verbatim, contents not normalized
last = wildcardMatcher.end();
}
if (last < value.length()) {
BytesRef normalized = normalizer.normalize(fieldname, value.substring(last));
sb.append(normalized); // (A) trailing gap, same gap handling
}
The in-code comment states the intent is to "normalize everything except wildcard characters." That intent is implemented only for the operator-preservation case. The two failure modes map directly onto the marked lines:
- Bug 1 is line (B). The escape match
\X is appended as-is, so the escaped character is never passed to normalizer.normalize(...).
- Bug 2 is line (A). A character that is not an ASCII
?/* falls into a normalized gap. The normalizer may emit *, ?, or \, and that output is appended without re-escaping, so emitted operators are interpreted as wildcards.
Steps to Reproduce
The pattern_replace char filter below mimics ICU NFKC for *, ?, and \ in isolation. Behavior is identical with the ICU plugin.
PUT /icu-wildcard-bug-test
{
"settings": {
"analysis": {
"normalizer": {
"icu_nfkc_mimic": {
"type": "custom",
"char_filter": ["fw_star", "fw_question", "fw_backslash"]
}
},
"char_filter": {
"fw_star": { "type": "pattern_replace", "pattern": "*", "replacement": "*" },
"fw_question": { "type": "pattern_replace", "pattern": "?", "replacement": "?" },
"fw_backslash":{ "type": "pattern_replace", "pattern": "\", "replacement": "\\\\" }
}
}
},
"mappings": {
"properties": {
"label": { "type": "keyword" },
"kw_vanilla": { "type": "keyword" },
"kw_normalized": { "type": "keyword", "normalizer": "icu_nfkc_mimic" }
}
}
}
POST /icu-wildcard-bug-test/_bulk
{ "index": { "_id": "1" } }
{ "label": "fw-star", "kw_vanilla": "東京*大阪", "kw_normalized": "東京*大阪" }
{ "index": { "_id": "2" } }
{ "label": "fw-question", "kw_vanilla": "東京?大阪", "kw_normalized": "東京?大阪" }
{ "index": { "_id": "3" } }
{ "label": "fw-backslash","kw_vanilla": "東京\大阪", "kw_normalized": "東京\大阪" }
{ "index": { "_id": "4" } }
{ "label": "no-special", "kw_vanilla": "東京大阪", "kw_normalized": "東京大阪" }
Index-time normalization is correct. Confirmed stored values: 東京*大阪 to 東京*大阪, 東京?大阪 to 東京?大阪, 東京\大阪 to 東京\大阪, 東京大阪 to 東京大阪.
Bug 1, escape contents bypassed. A user typing the fullwidth form and escaping it for a literal match gets zero hits:
GET /icu-wildcard-bug-test/_search
{ "query": { "wildcard": { "kw_normalized": { "value": "東京\\*大阪" } } } }
| Query value |
Field |
Expected |
Actual |
東京\*大阪 |
kw_normalized |
1 (fw-star) |
0 |
東京\*大阪 |
kw_normalized |
1 (fw-star) |
1 (only works if caller types the post-normalization ASCII form) |
東京\*大阪 |
kw_vanilla |
1 (fw-star) |
1 |
Bug 2, output not re-escaped. A user searching the fullwidth literal without an escape gets wildcard behavior:
GET /icu-wildcard-bug-test/_search
{ "query": { "wildcard": { "kw_normalized": { "value": "東京*大阪" } } } }
| Query value |
Expected |
Actual |
Matched |
東京*大阪 |
1 (fw-star, literal *) |
4 |
all docs |
東京?大阪 |
1 (fw-question, literal ?) |
3 |
fw-star, fw-question, fw-backslash |
東京*大阪 |
4 (operator *) |
4 |
all docs (confirms operator path is correct) |
Impact: no valid query formulation exists
| User intent |
Query |
Result |
Cause |
Literal * via fullwidth form |
東京*大阪 |
wrong, wildcard behavior |
Bug 2 |
Literal * via fullwidth escape |
東京\*大阪 |
wrong, 0 results |
Bug 1 |
Literal * via ASCII escape |
東京\*大阪 |
correct, but defeats normalization |
caller must know the post-normalization form |
* as operator |
東京* |
correct |
operator path is fine |
The only working formulation forces the caller to be normalizer aware and type the post-normalization ASCII form, which defeats the purpose of normalization on search terms. For fullwidth \ the failure is worse, since the stray \ generally yields zero matches rather than an over-inclusive set, so affected users silently get no results.
Expected behavior
The parser already segments the query into typed tokens (escapes and operator runs), so the infrastructure is in place. Two handling changes are needed inside normalizeWildcardPattern:
- For an escape match
\X: strip the \, normalize X as a literal, re-escape if the normalized output is *, ?, or \, then append. The \ is an instruction to treat the next character as literal, not data to preserve unmodified.
- For normalized gaps: after normalizing a chunk, escape any
*, ?, or \ that the normalizer produced, so emitted operators are treated as the literal data they came from.
Operator runs in group 2 should continue to be preserved as-is, which is the existing correct behavior.
Related issues
Logs (if relevant)
No response
Elasticsearch Version
Reproduced on
main. The relevant code is unchanged across recent 8.x and 9.x releases.Installed Plugins
analysis-icu(for the real-world trigger). The reproduction below uses a built-inpattern_replacechar filter and needs no plugins.Java Version
N/A (logic bug, version independent)
OS Version
N/A
Problem Description
When a
keywordfield uses a normalizer whose output adds, removes, or rewrites the wildcard control characters*,?, or\,wildcardqueries return incorrect results. The most common trigger is the ICU Analysis plugin's NFKC normalization, which maps fullwidth East Asian forms to ASCII (*to*,?to?,\to\). The same defect applies to any normalizer that maps in either direction across these characters.There are two distinct, direction-dependent failure modes:
\*, returns zero hits. The whole\Xunit is appended verbatim, so the escaped character is never normalized. The query then looks for a character that does not exist in the index.*, is normalized to*and then treated as a wildcard operator. A character that was data silently becomes an operator and matches far more documents than intended. With?it becomes a single-character wildcard; with fullwidth\the resulting stray\typically produces zero matches.Which bug you hit depends on the normalizer direction. A normalizer that maps
*to*exhibits Bug 1. A normalizer that maps*to*(ICU NFKC) exhibits Bug 2. As a result there is no normalizer that rewrites to or from these base forms for which literal wildcard search behaves correctly, and no query formulation that recovers correct behavior (see the impact table below).Root cause
The normalization-aware wildcard path is
StringFieldType.normalizeWildcardPattern(...), called from bothKeywordFieldMapperwildcard query paths. It segments the query with:Group 1 matches an escape sequence
\X. Group 2 matches runs of literal?/*. The loop normalizes the gaps between matches and appends each match verbatim:The in-code comment states the intent is to "normalize everything except wildcard characters." That intent is implemented only for the operator-preservation case. The two failure modes map directly onto the marked lines:
\Xis appended as-is, so the escaped character is never passed tonormalizer.normalize(...).?/*falls into a normalized gap. The normalizer may emit*,?, or\, and that output is appended without re-escaping, so emitted operators are interpreted as wildcards.Steps to Reproduce
The
pattern_replacechar filter below mimics ICU NFKC for*,?, and\in isolation. Behavior is identical with the ICU plugin.Index-time normalization is correct. Confirmed stored values:
東京*大阪to東京*大阪,東京?大阪to東京?大阪,東京\大阪to東京\大阪,東京大阪to東京大阪.Bug 1, escape contents bypassed. A user typing the fullwidth form and escaping it for a literal match gets zero hits:
東京\*大阪kw_normalizedfw-star)東京\*大阪kw_normalizedfw-star)東京\*大阪kw_vanillafw-star)Bug 2, output not re-escaped. A user searching the fullwidth literal without an escape gets wildcard behavior:
東京*大阪fw-star, literal*)東京?大阪fw-question, literal?)fw-star,fw-question,fw-backslash東京*大阪*)Impact: no valid query formulation exists
*via fullwidth form東京*大阪*via fullwidth escape東京\*大阪*via ASCII escape東京\*大阪*as operator東京*The only working formulation forces the caller to be normalizer aware and type the post-normalization ASCII form, which defeats the purpose of normalization on search terms. For fullwidth
\the failure is worse, since the stray\generally yields zero matches rather than an over-inclusive set, so affected users silently get no results.Expected behavior
The parser already segments the query into typed tokens (escapes and operator runs), so the infrastructure is in place. Two handling changes are needed inside
normalizeWildcardPattern:\X: strip the\, normalizeXas a literal, re-escape if the normalized output is*,?, or\, then append. The\is an instruction to treat the next character as literal, not data to preserve unmodified.*,?, or\that the normalizer produced, so emitted operators are treated as the literal data they came from.Operator runs in group 2 should continue to be preserved as-is, which is the existing correct behavior.
Related issues
Logs (if relevant)
No response