Repository navigation
[ESQL] Speed up REPLACE on constant regex - #149033
Conversation
When REPLACE is called with a constant regex, two cheap and correctness-preserving fast paths now skip the bulk of the work: 1. Anchored-prefix reject: if the compiled pattern begins with `^` or `\A` followed by literal characters, extract those bytes once and reject rows whose UTF-8 input doesn't start with them — before the per-row UTF-8 to UTF-16 conversion and Matcher allocation. The extractor is deliberately conservative: it bails on alternation, character classes, look-arounds, multiline / case-insensitive / comments / canonical-equivalence modes, and quantifiers that make the preceding literal optional. 2. Dictionary memoization on OrdinalBytesRefBlock: when both regex and replacement are constants and the input column is dictionary-encoded (typical for keyword fields backed by Lucene doc values), apply REPLACE once per dictionary entry and emit a new ordinal block that reuses the input ordinals. For a column with N rows backed by D dictionary entries this drops regex work from N calls to D. The path is gated on a density check and single-valued input; any dictionary-entry overflow falls back to the per-row code so warning semantics are preserved exactly. The two paths compose: the prefix check applies inside both the per-row loop and the dictionary loop. Developed using AI-assisted tooling
|
Pinging @elastic/es-analytical-engine (Team:Analytics) |
|
Hi @costin, I've created a changelog YAML for you. |
🔍 Preview links for changed docs⏳ Building and deploying preview... View progress This comment will be updated with preview links when the build is complete. |
ℹ️ Important: Docs version tagging👋 Thanks for updating the docs! Just a friendly reminder that our docs are now cumulative. This means all 9.x versions are documented on the same page and published off of the main branch, instead of creating separate pages for each minor version. We use applies_to tags to mark version-specific features and changes. Expand for a quick overviewWhen to use applies_to tags:✅ At the page level to indicate which products/deployments the content applies to (mandatory) What NOT to do:❌ Don't remove or replace information that applies to an older version 🤔 Need help?
|
* ESQL: REPLACE fast-path review fixes Three small follow-ups to #149033: - Rename containsTopLevelAlternation to containsUnquotedAlternation. The method bails on any unescaped `|`, including ones nested inside groups (e.g. `^foo(a|b)`). The old name suggested only top-level alternation was rejected, which misled future readers. - Account for the constant literalPrefix byte[] and newStr BytesRef in ReplaceConstantOrdinalEvaluator.baseRamBytesUsed. Both are usually small, but they were the only constants we held without RAM accounting. - Add a regression test asserting `\A` and `^` produce identical prefix extraction across quantifier shapes (`?`, `+`, `{0,3}`). Developed using AI-assisted tooling * Update docs/changelog/149167.yaml
|
Retroactively supersedes elastic/esql-planning#700 (URL-domain REPLACE byte-scan) — different mechanism, broader coverage. |
Summary
Two complementary, correctness-preserving fast paths for
REPLACE(field, regex, newStr)when the regex is a constant:^or\Afollowed by literal characters, those bytes are extracted once at plan time. Per row, a byte-level prefix check on the UTF-8 input rejects non-matching rows before paying for UTF-8 → UTF-16 conversion and Matcher allocation. The extractor is deliberately conservative: it bails on alternation, character classes, look-arounds,(?m)/(?i)/(?x)/CANON_EQ, and quantifiers that make the preceding literal optional.REPLACEis applied once per dictionary entry instead of once per row. The result re-uses the input's ordinal column with a transformed dictionary. For a column with N rows backed by D dictionary entries, this drops regex work from N calls to D — typically a 10–50x reduction.The dictionary path is gated on
OrdinalBytesRefBlock#isDense()and a single-valued check, and falls back to the existing per-row evaluator if any dictionary entry hits the result-size limit — so observable warning semantics are unchanged.Test plan
ReplaceTestsandReplaceStaticTestsstill pass.ReplaceStaticTestscases cover the prefix extractor across a wide range of regex constructs (anchors, escapes, quantifiers,\Q...\E, flag combinations).ReplaceOrdinalTestsexercises the dictionary path on dense + null + sparse + overflow scenarios.:x-pack:plugin:esql:precommitpasses.Developed with AI-assisted tooling