Skip to content

Update to lucene 10.5 - #151959

Merged
jimczi merged 222 commits into
mainfrom
lucene_10_5_snapshot
Jun 30, 2026
Merged

jimczi merged 222 commits into
mainfrom
lucene_10_5_snapshot

Conversation

@romseygeek

Copy link
Copy Markdown
Contributor

No description provided.

benwtrent and others added 30 commits March 11, 2026 07:56
@elasticsearchmachine

Copy link
Copy Markdown
Collaborator

Pinging @elastic/es-storage-engine (Team:StorageEngine)

@elasticsearchmachine

Copy link
Copy Markdown
Collaborator

Hi @romseygeek, I've created a changelog YAML for you.

@coderabbitai

coderabbitai Bot commented Jun 29, 2026

Copy link
Copy Markdown
Contributor

Caution

Review failed

An error occurred during the review process. Please try again later.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch lucene_10_5_snapshot

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

1 similar comment
@coderabbitai

coderabbitai Bot commented Jun 29, 2026

Copy link
Copy Markdown
Contributor

Caution

Review failed

An error occurred during the review process. Please try again later.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch lucene_10_5_snapshot

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@iverase

iverase commented Jun 29, 2026

Copy link
Copy Markdown
Contributor

Buildkite benchmark this with noaa-1n-1g

@infra-vault-gh-plugin-prod

infra-vault-gh-plugin-prod Bot commented Jun 29, 2026 •

Copy link
Copy Markdown

💚 Build Succeeded

This build ran two noaa-1n-1g benchmarks to evaluate performance impact of this PR.

History

cc @romseygeek

jimczi added 10 commits June 29, 2026 12:26
getRandomVectorScorerSupplierForMerge built the index vectors from the
whole segment data file instead of slicing to the field region, so the
HNSW graph was scored against bytes at the wrong file offset (silent
recall degradation on BBQ-HNSW merges). Read them via
OffHeapBinarizedVectorValues.load(...), mirroring the other read paths.
readDocValueSkipperMeta guarded the maxValueCount read on VERSION_CURRENT
instead of VERSION_SKIPPER_MAX_VALUE_COUNT. Identical today, but a future
version bump would make v1 segments skip a field they did write,
desyncing the metadata stream. Matches AbstractTSDBDocValuesProducer.
getRandomVectorScorerSupplierForMerge closed the score input on the error
path after openInput but left the temp queries file on disk. Delete it,
matching the earlier write-failure catch.
writeBinarizedQueryData dropped the [0, 0xffff] bounds assert before
writeShort when it moved from the writer. The flat write paths keep it;
restore it for parity.
The fork existed only to backport apache/lucene#15723 (the
tryPopulateCache hook for intra-segment cache population). Lucene 10.5
ships that hook, so ElasticsearchLRUQueryCache extends the upstream
LRUQueryCache directly and the fork is deleted.
Adds coverage for the intra-segment cache gate doc-partitioning relies on:
a query whose weight is an OptionalCachingWeight populates the cache only
when startCaching grants the claim, and falls back to an uncached scorer
otherwise. Guards the seam now served by Lucene 10.5 (apache/lucene#15723).
Clarify that mergeOneField invokes the per-field beginIvfFieldMerge hook
for every field (including non-float encodings) for subclasses such as
ESNextDiskBBQVectorsWriter, and that the resolved config is intentionally
unused in the base class. Drops the dead local to match the comment.
@jimczi
jimczi removed the request for review from a team June 29, 2026 16:30

@jimczi jimczi left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@jimczi
jimczi merged commit 09bd9f5 into main Jun 30, 2026
39 of 40 checks passed
@jimczi
jimczi deleted the lucene_10_5_snapshot branch June 30, 2026 11:24
smalyshev added a commit to smalyshev/elasticsearch that referenced this pull request Jul 1, 2026
Upgrade to Lucene 10.5

Co-authored-by: Benjamin Trent <4357155+benwtrent@users.noreply.github.com>
Co-authored-by: Ignacio Vera <ivera@apache.org>
Co-authored-by: elasticsearchmachine <infra-root+elasticsearchmachine@elastic.co>
Co-authored-by: John Wagster <john.wagster@elastic.co>
Co-authored-by: Luigi Dell'Aquila <luigi.dellaquila@gmail.com>
Co-authored-by: Panagiotis Bailis <pmpailis@gmail.com>
Co-authored-by: Mouhcine Aitounejjar <5037925+mouhc1ne@users.noreply.github.com>
Co-authored-by: Stanislav Malyshev <stas.malyshev@elastic.co>
Co-authored-by: Simon Cooper <simon.cooper@elastic.co>
Co-authored-by: Ignacio Vera <ignacio.vera@elastic.co>
Co-authored-by: Mayya Sharipova <mayya.sharipova@elastic.co>
Co-authored-by: Tommaso Teofili <tommaso.teofili@elastic.co>
Co-authored-by: Salvatore Campagna <93581129+salvatore-campagna@users.noreply.github.com>
Co-authored-by: Alexander Reelsen <alexander.reelsen@elastic.co>
Co-authored-by: Ben Chaplin <benchaplin@protonmail.com>
Co-authored-by: Jim Ferenczi <jim.ferenczi@elastic.co>
elasticsearchmachine pushed a commit that referenced this pull request Jul 6, 2026
PR #151959 upgraded Lucene to 10.5 and broke a couple of things in the
GPU codec: first the tests (the byte RAM estimate test, now skipped),
and then the byte merge path started logging "Cannot get merged raw
vectors from scorer. Performances will be degraded." and falling back to
a slow re-merge on every merge.

Lucene 10.5 changed how merging works: the merged raw/quantized vectors
are now written to the flat file first (mergeOneFlatVectorField), and
the graph is built afterwards. Previously the vectors weren't
materialized yet, so the writer had to reach them through a scorer
supplier and extract the underlying values by reflection - and after the
upgrade that supplier became a generic wrapper the reflection no longer
recognized, hence the warning and slow path.

Now that the merged vectors already live in the flat file before the
graph build, read them straight from the flat reader: its values expose
the mmap-backed file slice via HasIndexSlice (both FloatVectorValues and
the quantized BaseQuantizedByteVectorValues), so there is no need for
the scorer supplier or reflection. Drop the now-unused
VectorsFormatReflectionUtils.
burqen pushed a commit to burqen/elasticsearch that referenced this pull request Jul 7, 2026
PR elastic#151959 upgraded Lucene to 10.5 and broke a couple of things in the
GPU codec: first the tests (the byte RAM estimate test, now skipped),
and then the byte merge path started logging "Cannot get merged raw
vectors from scorer. Performances will be degraded." and falling back to
a slow re-merge on every merge.

Lucene 10.5 changed how merging works: the merged raw/quantized vectors
are now written to the flat file first (mergeOneFlatVectorField), and
the graph is built afterwards. Previously the vectors weren't
materialized yet, so the writer had to reach them through a scorer
supplier and extract the underlying values by reflection - and after the
upgrade that supplier became a generic wrapper the reflection no longer
recognized, hence the warning and slow path.

Now that the merged vectors already live in the flat file before the
graph build, read them straight from the flat reader: its values expose
the mmap-backed file slice via HasIndexSlice (both FloatVectorValues and
the quantized BaseQuantizedByteVectorValues), so there is no need for
the scorer supplier or reflection. Drop the now-unused
VectorsFormatReflectionUtils.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.