Skip to content

[ESQL] Speed up REPLACE on constant regex - #149033

Merged
costin merged 2 commits into
elastic:mainfrom
costin:esql/replace-prefix-fastpath
May 15, 2026
Merged

costin merged 2 commits into
elastic:mainfrom
costin:esql/replace-prefix-fastpath

Conversation

@costin

@costin costin commented May 13, 2026

Copy link
Copy Markdown
Member

Summary

Two complementary, correctness-preserving fast paths for REPLACE(field, regex, newStr) when the regex is a constant:

  1. Anchored-prefix reject — when the compiled pattern starts with ^ or \A followed by literal characters, those bytes are extracted once at plan time. Per row, a byte-level prefix check on the UTF-8 input rejects non-matching rows before paying for UTF-8 → UTF-16 conversion and Matcher allocation. The extractor is deliberately conservative: it bails on alternation, character classes, look-arounds, (?m) / (?i) / (?x) / CANON_EQ, and quantifiers that make the preceding literal optional.
  2. Dictionary memoization on dictionary-encoded inputs — when both the regex and the replacement are constants and the input column is dictionary-encoded (typical for keyword fields backed by Lucene doc values), REPLACE is applied once per dictionary entry instead of once per row. The result re-uses the input's ordinal column with a transformed dictionary. For a column with N rows backed by D dictionary entries, this drops regex work from N calls to D — typically a 10–50x reduction.

The dictionary path is gated on OrdinalBytesRefBlock#isDense() and a single-valued check, and falls back to the existing per-row evaluator if any dictionary entry hits the result-size limit — so observable warning semantics are unchanged.

Test plan

  • Existing ReplaceTests and ReplaceStaticTests still pass.
  • New ReplaceStaticTests cases cover the prefix extractor across a wide range of regex constructs (anchors, escapes, quantifiers, \Q...\E, flag combinations).
  • New ReplaceOrdinalTests exercises the dictionary path on dense + null + sparse + overflow scenarios.
  • :x-pack:plugin:esql:precommit passes.

Developed with AI-assisted tooling

When REPLACE is called with a constant regex, two cheap and
correctness-preserving fast paths now skip the bulk of the work:

1. Anchored-prefix reject: if the compiled pattern begins with `^`
   or `\A` followed by literal characters, extract those bytes once
   and reject rows whose UTF-8 input doesn't start with them — before
   the per-row UTF-8 to UTF-16 conversion and Matcher allocation. The
   extractor is deliberately conservative: it bails on alternation,
   character classes, look-arounds, multiline / case-insensitive /
   comments / canonical-equivalence modes, and quantifiers that make
   the preceding literal optional.

2. Dictionary memoization on OrdinalBytesRefBlock: when both regex and
   replacement are constants and the input column is dictionary-encoded
   (typical for keyword fields backed by Lucene doc values), apply
   REPLACE once per dictionary entry and emit a new ordinal block that
   reuses the input ordinals. For a column with N rows backed by D
   dictionary entries this drops regex work from N calls to D. The
   path is gated on a density check and single-valued input; any
   dictionary-entry overflow falls back to the per-row code so warning
   semantics are preserved exactly.

The two paths compose: the prefix check applies inside both the per-row
loop and the dictionary loop.

Developed using AI-assisted tooling
@costin
costin requested a review from bpintea May 13, 2026 23:37
@costin
costin enabled auto-merge (squash) May 13, 2026 23:37
@elasticsearchmachine elasticsearchmachine added Team:Analytics Meta label for analytical engine team (ESQL/Aggs/Geo) v9.5.0 labels May 13, 2026
@elasticsearchmachine

Copy link
Copy Markdown
Collaborator

Pinging @elastic/es-analytical-engine (Team:Analytics)

@elasticsearchmachine

Copy link
Copy Markdown
Collaborator

Hi @costin, I've created a changelog YAML for you.

@github-actions

Copy link
Copy Markdown
Contributor

🔍 Preview links for changed docs

⏳ Building and deploying preview... View progress

This comment will be updated with preview links when the build is complete.

@github-actions

Copy link
Copy Markdown
Contributor

ℹ️ Important: Docs version tagging

👋 Thanks for updating the docs! Just a friendly reminder that our docs are now cumulative. This means all 9.x versions are documented on the same page and published off of the main branch, instead of creating separate pages for each minor version.

We use applies_to tags to mark version-specific features and changes.

Expand for a quick overview

When to use applies_to tags:

✅ At the page level to indicate which products/deployments the content applies to (mandatory)
✅ When features change state (e.g. preview, ga) in a specific version
✅ When availability differs across deployments and environments

What NOT to do:

❌ Don't remove or replace information that applies to an older version
❌ Don't add new information that applies to a specific version without an applies_to tag
❌ Don't forget that applies_to tags can be used at the page, section, and inline level

🤔 Need help?

@bpintea bpintea left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LG, review out-of-band.

@costin
costin merged commit 89bf32a into elastic:main May 15, 2026
37 checks passed
@costin
costin deleted the esql/replace-prefix-fastpath branch May 15, 2026 10:35
costin added a commit that referenced this pull request May 15, 2026
* ESQL: REPLACE fast-path review fixes

Three small follow-ups to #149033:

- Rename containsTopLevelAlternation to containsUnquotedAlternation.
  The method bails on any unescaped `|`, including ones nested inside
  groups (e.g. `^foo(a|b)`). The old name suggested only top-level
  alternation was rejected, which misled future readers.
- Account for the constant literalPrefix byte[] and newStr BytesRef in
  ReplaceConstantOrdinalEvaluator.baseRamBytesUsed. Both are usually
  small, but they were the only constants we held without RAM
  accounting.
- Add a regression test asserting `\A` and `^` produce identical prefix
  extraction across quantifier shapes (`?`, `+`, `{0,3}`).

Developed using AI-assisted tooling

* Update docs/changelog/149167.yaml
@quackaplop

Copy link
Copy Markdown
Member

Retroactively supersedes elastic/esql-planning#700 (URL-domain REPLACE byte-scan) — different mechanism, broader coverage.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

:Analytics/ES|QL AKA ESQL >enhancement ES|QL|DS ES|QL datasources Team:Analytics Meta label for analytical engine team (ESQL/Aggs/Geo) v9.4.0 v9.5.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants