Skip to content

Simplify field matching in exclude source vectors - #156466

Merged
elasticsearchmachine merged 4 commits into
elastic:mainfrom
mayya-sharipova:fix/simplify-exclude-source-vectors-matching
Aug 12, 2026
Merged

elasticsearchmachine merged 4 commits into
elastic:mainfrom
mayya-sharipova:fix/simplify-exclude-source-vectors-matching

Conversation

@mayya-sharipova

Copy link
Copy Markdown
Contributor

The fetch phase decides which vector fields to strip from _source by
compiling the field patterns of the request into an automaton. It did so
unconditionally, before establishing whether the mapping held any vector
embeddings at all, so an index with none paid for a matcher that had
nothing to match. Worse, a request asking for enough field patterns
exceeded Lucene's determinization work limit, and the resulting
TooComplexToDeterminizeException failed the search even on indices where
not a single field would have been excluded.

Resolve the vector fields first and return early when the mapping has
none, which skips the whole step for such indices.

Then match the remaining patterns one at a time with Regex#simpleMatch
rather than compiling them. Compiling amortizes its cost over the
strings tested against the result, and the only strings tested here are
the vector fields, of which a mapping holds a handful, while a request
can carry arbitrarily many patterns. The build therefore dominated and
was never repaid. Matching individually suits this shape better and,
because nothing is compiled, no work limit applies.

The inference field patterns come from the mapping rather than from the
request, so their number is bounded. They keep a compiled matcher, now
built through Regex#simpleMatcher so that the automaton is skipped when
it is not needed.

The fetch phase decides which vector fields to strip from _source by
compiling the field patterns of the request into an automaton. It did so
unconditionally, before establishing whether the mapping held any vector
embeddings at all, so an index with none paid for a matcher that had
nothing to match. Worse, a request asking for enough field patterns
exceeded Lucene's determinization work limit, and the resulting
TooComplexToDeterminizeException failed the search even on indices where
not a single field would have been excluded.

Resolve the vector fields first and return early when the mapping has
none, which skips the whole step for such indices.

Then match the remaining patterns one at a time with Regex#simpleMatch
rather than compiling them. Compiling amortizes its cost over the
strings tested against the result, and the only strings tested here are
the vector fields, of which a mapping holds a handful, while a request
can carry arbitrarily many patterns. The build therefore dominated and
was never repaid. Matching individually suits this shape better and,
because nothing is compiled, no work limit applies.

The inference field patterns come from the mapping rather than from the
request, so their number is bounded. They keep a compiled matcher, now
built through Regex#simpleMatcher so that the automaton is skipped when
it is not needed.
@elasticsearchmachine elasticsearchmachine added v9.6.0 needs:triage Requires assignment of a team area label labels Aug 11, 2026
@mayya-sharipova mayya-sharipova added >bug :Search Relevance/Vectors Vector search auto-backport Automatically create backport pull requests when merged v9.5.1 labels Aug 11, 2026
@elasticsearchmachine elasticsearchmachine added Team:Search Relevance Meta label for the Search Relevance team in Elasticsearch and removed needs:triage Requires assignment of a team area label labels Aug 11, 2026
@elasticsearchmachine

Copy link
Copy Markdown
Collaborator

Pinging @elastic/es-search-relevance (Team:Search Relevance)

@elasticsearchmachine

Copy link
Copy Markdown
Collaborator

Hi @mayya-sharipova, I've created a changelog YAML for you.

@github-actions

github-actions Bot commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

🔍 Preview links for changed docs

⏳ Building and deploying preview... View progress

This comment will be updated with preview links when the build is complete.

@github-actions

Copy link
Copy Markdown
Contributor

ℹ️ Important: Docs version tagging

👋 Thanks for updating the docs! Just a friendly reminder that our docs are now cumulative. This means all 9.x versions are documented on the same page and published off of the main branch, instead of creating separate pages for each minor version.

We use applies_to tags to mark version-specific features and changes.

Expand for a quick overview

When to use applies_to tags:

✅ At the page level to indicate which products/deployments the content applies to (mandatory)
✅ When features change state (e.g. preview, ga) in a specific version
✅ When availability differs across deployments and environments

What NOT to do:

❌ Don't remove or replace information that applies to an older version
❌ Don't add new information that applies to a specific version without an applies_to tag
❌ Don't forget that applies_to tags can be used at the page, section, and inline level

🤔 Need help?

@mayya-sharipova
mayya-sharipova requested a review from jimczi August 11, 2026 18:42

@jimczi jimczi left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice find. The approach looks right to me: not compiling the request patterns is what the rest of the fetch path already does. FieldTypeLookup#getMatchingFieldNames returns the keySet for a match-all, a hash lookup for a concrete name, and only falls back to per pattern Regex#simpleMatch for wildcards.

Comments inline. One thing outside the diff: labels are v9.5.1 and v9.6.0, but this shipped in 9.2.0 and the open patch versions are v9.4.6, v9.5.2 and v9.6.0. Looks like v9.5.1 is already cut and v9.4.6 is missing.

Comment thread server/src/main/java/org/elasticsearch/index/get/ShardGetService.java Outdated
Comment thread docs/changelog/156466.yaml Outdated
Precompute the vector embedding fields on MappingLookup rather than
streaming every field type on each fetch. The mapping is immutable, so
the candidate set is established once, in the constructor loop that
already builds inferenceFields and syntheticVectorFields, and the early
return becomes an isEmpty check with no per request allocation. That is
worth it because the requests reaching this path are the wide ones.

The existing syntheticVectorFields cannot be reused for this. It is
keyed on syntheticVectorsLoader, which is null unless the mapper carries
excludeSourceVectors, so the set is empty when
index.mapping.exclude_source_vectors is false. This method is still
reached in that case through a request level exclude_vectors, so the
candidates have to be resolved independently of the setting.

In the test fixture, ask for a plain match-all alongside the concrete
names instead of a crafted trailing wildcard. That is what a client
sends, it reaches the determinization limit with fewer patterns, and it
matches the vector field, so the late exclude branch is covered rather
than only the up front one. Assert the precondition as well, since the
fixture covers nothing if that limit ever moves.

Concrete names alone never exceed the limit at any count: a string union
is already minimal and determinize short circuits on it. It is the
match-all that forces the subset construction, so document that as the
mechanism instead.

Also use Strings#format, which forbiddenApisTest requires over
String#format, and reword the changelog to name the symptom.
@mayya-sharipova

mayya-sharipova commented Aug 12, 2026 •

Copy link
Copy Markdown
Contributor Author

@jimczi Thanks Jim for the review, all addressed. Please continue the review.

@jimczi jimczi left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks Mayya, all addressed. Precomputing the set on MappingLookup came out cleaner than what I suggested, and the javadoc on why it is not syntheticVectorFields is worth having.

Ran the unit tests and the search.vectors / get/100_synthetic_source yaml suites locally, all green. The 9.4.5 bwc failure is the MixedClusterEsqlSpecLookupJoinIT wave that is failing on main, unrelated.

LGTM.

@mayya-sharipova mayya-sharipova added the auto-merge-without-approval Automatically merge pull request when CI checks pass (NB doesn't wait for reviews!) label Aug 12, 2026
@elasticsearchmachine
elasticsearchmachine merged commit 8528446 into elastic:main Aug 12, 2026
37 checks passed
@mayya-sharipova
mayya-sharipova deleted the fix/simplify-exclude-source-vectors-matching branch August 12, 2026 20:29
@elasticsearchmachine

Copy link
Copy Markdown
Collaborator

💔 Backport failed

Status Branch Result
❌ 9.4 Commit could not be cherrypicked due to conflicts
✅ 9.5

You can use sqren/backport to manually backport by running backport --upstream elastic/elasticsearch --pr 156466

elasticsearchmachine pushed a commit that referenced this pull request Aug 12, 2026
The fetch phase decides which vector fields to strip from _source by
compiling the field patterns of the request into an automaton. It did so
unconditionally, before establishing whether the mapping held any vector
embeddings at all, so an index with none paid for a matcher that had
nothing to match. Worse, a request asking for enough field patterns
exceeded Lucene's determinization work limit, and the resulting
TooComplexToDeterminizeException failed the search even on indices where
not a single field would have been excluded.

Resolve the vector fields first and return early when the mapping has
none, which skips the whole step for such indices.

Then match the remaining patterns one at a time with Regex#simpleMatch
rather than compiling them. Compiling amortizes its cost over the
strings tested against the result, and the only strings tested here are
the vector fields, of which a mapping holds a handful, while a request
can carry arbitrarily many patterns. The build therefore dominated and
was never repaid. Matching individually suits this shape better and,
because nothing is compiled, no work limit applies.

The inference field patterns come from the mapping rather than from the
request, so their number is bounded. They keep a compiled matcher, now
built through Regex#simpleMatcher so that the automaton is skipped when
it is not needed.
elasticsearchmachine pushed a commit that referenced this pull request Aug 12, 2026
…56613)

* Simplify field matching in exclude source vectors (#156466)

The fetch phase decides which vector fields to strip from _source by
compiling the field patterns of the request into an automaton. It did so
unconditionally, before establishing whether the mapping held any vector
embeddings at all, so an index with none paid for a matcher that had
nothing to match. Worse, a request asking for enough field patterns
exceeded Lucene's determinization work limit, and the resulting
TooComplexToDeterminizeException failed the search even on indices where
not a single field would have been excluded.

Resolve the vector fields first and return early when the mapping has
none, which skips the whole step for such indices.

Then match the remaining patterns one at a time with Regex#simpleMatch
rather than compiling them. Compiling amortizes its cost over the
strings tested against the result, and the only strings tested here are
the vector fields, of which a mapping holds a handful, while a request
can carry arbitrarily many patterns. The build therefore dominated and
was never repaid. Matching individually suits this shape better and,
because nothing is compiled, no work limit applies.

The inference field patterns come from the mapping rather than from the
request, so their number is bounded. They keep a compiled matcher, now
built through Regex#simpleMatcher so that the automaton is skipped when
it is not needed.

(cherry picked from commit 8528446)

* Retrigger CI after gradle.org 503
ChrisHegarty pushed a commit to ChrisHegarty/elasticsearch that referenced this pull request Aug 13, 2026
The fetch phase decides which vector fields to strip from _source by
compiling the field patterns of the request into an automaton. It did so
unconditionally, before establishing whether the mapping held any vector
embeddings at all, so an index with none paid for a matcher that had
nothing to match. Worse, a request asking for enough field patterns
exceeded Lucene's determinization work limit, and the resulting
TooComplexToDeterminizeException failed the search even on indices where
not a single field would have been excluded.

Resolve the vector fields first and return early when the mapping has
none, which skips the whole step for such indices.

Then match the remaining patterns one at a time with Regex#simpleMatch
rather than compiling them. Compiling amortizes its cost over the
strings tested against the result, and the only strings tested here are
the vector fields, of which a mapping holds a handful, while a request
can carry arbitrarily many patterns. The build therefore dominated and
was never repaid. Matching individually suits this shape better and,
because nothing is compiled, no work limit applies.

The inference field patterns come from the mapping rather than from the
request, so their number is bounded. They keep a compiled matcher, now
built through Regex#simpleMatcher so that the automaton is skipped when
it is not needed.
ncordon pushed a commit to ncordon/elasticsearch that referenced this pull request Aug 24, 2026
The fetch phase decides which vector fields to strip from _source by
compiling the field patterns of the request into an automaton. It did so
unconditionally, before establishing whether the mapping held any vector
embeddings at all, so an index with none paid for a matcher that had
nothing to match. Worse, a request asking for enough field patterns
exceeded Lucene's determinization work limit, and the resulting
TooComplexToDeterminizeException failed the search even on indices where
not a single field would have been excluded.

Resolve the vector fields first and return early when the mapping has
none, which skips the whole step for such indices.

Then match the remaining patterns one at a time with Regex#simpleMatch
rather than compiling them. Compiling amortizes its cost over the
strings tested against the result, and the only strings tested here are
the vector fields, of which a mapping holds a handful, while a request
can carry arbitrarily many patterns. The build therefore dominated and
was never repaid. Matching individually suits this shape better and,
because nothing is compiled, no work limit applies.

The inference field patterns come from the mapping rather than from the
request, so their number is bounded. They keep a compiled matcher, now
built through Regex#simpleMatcher so that the automaton is skipped when
it is not needed.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

auto-backport Automatically create backport pull requests when merged auto-merge-without-approval Automatically merge pull request when CI checks pass (NB doesn't wait for reviews!) backport pending >bug :Search Relevance/Vectors Vector search Team:Search Relevance Meta label for the Search Relevance team in Elasticsearch v9.4.6 v9.5.2 v9.6.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants