Repository navigation
Update to lucene 10.5 - #151959
Update to lucene 10.5#151959
Conversation
…to lucene_snapshot
|
Pinging @elastic/es-storage-engine (Team:StorageEngine) |
|
Hi @romseygeek, I've created a changelog YAML for you. |
|
Caution Review failedAn error occurred during the review process. Please try again later. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
1 similar comment
|
Caution Review failedAn error occurred during the review process. Please try again later. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Buildkite benchmark this with noaa-1n-1g |
💚 Build Succeeded
This build ran two noaa-1n-1g benchmarks to evaluate performance impact of this PR. History
cc @romseygeek |
getRandomVectorScorerSupplierForMerge built the index vectors from the whole segment data file instead of slicing to the field region, so the HNSW graph was scored against bytes at the wrong file offset (silent recall degradation on BBQ-HNSW merges). Read them via OffHeapBinarizedVectorValues.load(...), mirroring the other read paths.
readDocValueSkipperMeta guarded the maxValueCount read on VERSION_CURRENT instead of VERSION_SKIPPER_MAX_VALUE_COUNT. Identical today, but a future version bump would make v1 segments skip a field they did write, desyncing the metadata stream. Matches AbstractTSDBDocValuesProducer.
getRandomVectorScorerSupplierForMerge closed the score input on the error path after openInput but left the temp queries file on disk. Delete it, matching the earlier write-failure catch.
writeBinarizedQueryData dropped the [0, 0xffff] bounds assert before writeShort when it moved from the writer. The flat write paths keep it; restore it for parity.
The fork existed only to backport apache/lucene#15723 (the tryPopulateCache hook for intra-segment cache population). Lucene 10.5 ships that hook, so ElasticsearchLRUQueryCache extends the upstream LRUQueryCache directly and the fork is deleted.
Adds coverage for the intra-segment cache gate doc-partitioning relies on: a query whose weight is an OptionalCachingWeight populates the cache only when startCaching grants the claim, and falls back to an uncached scorer otherwise. Guards the seam now served by Lucene 10.5 (apache/lucene#15723).
Clarify that mergeOneField invokes the per-field beginIvfFieldMerge hook for every field (including non-float encodings) for subclasses such as ESNextDiskBBQVectorsWriter, and that the resolved config is intentionally unused in the base class. Drops the dead local to match the comment.
Upgrade to Lucene 10.5 Co-authored-by: Benjamin Trent <4357155+benwtrent@users.noreply.github.com> Co-authored-by: Ignacio Vera <ivera@apache.org> Co-authored-by: elasticsearchmachine <infra-root+elasticsearchmachine@elastic.co> Co-authored-by: John Wagster <john.wagster@elastic.co> Co-authored-by: Luigi Dell'Aquila <luigi.dellaquila@gmail.com> Co-authored-by: Panagiotis Bailis <pmpailis@gmail.com> Co-authored-by: Mouhcine Aitounejjar <5037925+mouhc1ne@users.noreply.github.com> Co-authored-by: Stanislav Malyshev <stas.malyshev@elastic.co> Co-authored-by: Simon Cooper <simon.cooper@elastic.co> Co-authored-by: Ignacio Vera <ignacio.vera@elastic.co> Co-authored-by: Mayya Sharipova <mayya.sharipova@elastic.co> Co-authored-by: Tommaso Teofili <tommaso.teofili@elastic.co> Co-authored-by: Salvatore Campagna <93581129+salvatore-campagna@users.noreply.github.com> Co-authored-by: Alexander Reelsen <alexander.reelsen@elastic.co> Co-authored-by: Ben Chaplin <benchaplin@protonmail.com> Co-authored-by: Jim Ferenczi <jim.ferenczi@elastic.co>
PR #151959 upgraded Lucene to 10.5 and broke a couple of things in the GPU codec: first the tests (the byte RAM estimate test, now skipped), and then the byte merge path started logging "Cannot get merged raw vectors from scorer. Performances will be degraded." and falling back to a slow re-merge on every merge. Lucene 10.5 changed how merging works: the merged raw/quantized vectors are now written to the flat file first (mergeOneFlatVectorField), and the graph is built afterwards. Previously the vectors weren't materialized yet, so the writer had to reach them through a scorer supplier and extract the underlying values by reflection - and after the upgrade that supplier became a generic wrapper the reflection no longer recognized, hence the warning and slow path. Now that the merged vectors already live in the flat file before the graph build, read them straight from the flat reader: its values expose the mmap-backed file slice via HasIndexSlice (both FloatVectorValues and the quantized BaseQuantizedByteVectorValues), so there is no need for the scorer supplier or reflection. Drop the now-unused VectorsFormatReflectionUtils.
PR elastic#151959 upgraded Lucene to 10.5 and broke a couple of things in the GPU codec: first the tests (the byte RAM estimate test, now skipped), and then the byte merge path started logging "Cannot get merged raw vectors from scorer. Performances will be degraded." and falling back to a slow re-merge on every merge. Lucene 10.5 changed how merging works: the merged raw/quantized vectors are now written to the flat file first (mergeOneFlatVectorField), and the graph is built afterwards. Previously the vectors weren't materialized yet, so the writer had to reach them through a scorer supplier and extract the underlying values by reflection - and after the upgrade that supplier became a generic wrapper the reflection no longer recognized, hence the warning and slow path. Now that the merged vectors already live in the flat file before the graph build, read them straight from the flat reader: its values expose the mmap-backed file slice via HasIndexSlice (both FloatVectorValues and the quantized BaseQuantizedByteVectorValues), so there is no need for the scorer supplier or reflection. Drop the now-unused VectorsFormatReflectionUtils.
No description provided.