Repository navigation
Advise MADV_RANDOM on blob cache regions backing vector data files - #150066
Conversation
…kernel read-ahead
Remove the separate StatelessKnnIndexTester class and the x-pack/qa/vector subproject. The stateless directory type is now registered directly in KnnIndexTester static initializer, alongside default and frozen. The stateless directory factory and cache stats logger are invoked via reflection (same pattern as the existing frozen/searchable-snapshots directory) since the stateless test artifact lives in an unnamed module.
…ix graph stats - Add skip_recall config option to skip brute-force exact NN computation for QPS-only benchmark runs - Fix race condition in KnnIndexer where ID_FIELD could get out of sync with the actual vector read from the dataset file by moving ordinal assignment into IndexVectorReader - Fix logGraphStats to unwrap PerFieldKnnVectorsFormat.FieldsReader before checking for HnswGraphProvider Co-authored-by: Cursor <cursoragent@cursor.com>
|
Pinging @elastic/es-search-relevance (Team:Search Relevance) |
|
Hi @ChrisHegarty, I've created a changelog YAML for you. |
🔍 Preview links for changed docs⏳ Building and deploying preview... View progress This comment will be updated with preview links when the build is complete. |
ℹ️ Important: Docs version tagging👋 Thanks for updating the docs! Just a friendly reminder that our docs are now cumulative. This means all 9.x versions are documented on the same page and published off of the main branch, instead of creating separate pages for each minor version. We use applies_to tags to mark version-specific features and changes. Expand for a quick overviewWhen to use applies_to tags:✅ At the page level to indicate which products/deployments the content applies to (mandatory) What NOT to do:❌ Don't remove or replace information that applies to an older version 🤔 Need help?
|
|
For non-compound (top-level) files, each file exclusively owns its blob — there are no shared regions. The exclusive range is [0, Long.MAX_VALUE), so every region gets MADV_RANDOM. No boundary fallback is needed. For compound files (.cfs), multiple sub-files are packed into a single blob and share cache regions at the boundaries. The table below shows how significant (or insignificant) those boundary regions are relative to the total .vec file size, using the default blob cache region size of 16 MiB:
At most 2 boundary regions (32 MiB) per sub-file fall back to MADV_NORMAL — one at each end where the sub-file shares a cache region with an adjacent file in the compound blob. For any non-trivial dataset (1M+ vectors), this is well under 5% of regions. For typical production workloads (millions of vectors, 768+ dims), over 99% of regions receive MADV_RANDOM. The small dataset case (100K vectors, 384 dims) is the worst case at 66.7% boundary — but at only 36.6 MiB total, the entire file fits comfortably in cache and read-ahead suppression matters far less. |
…madvise-random-vec-blob-cache
…nto madvise-random-vec-blob-cache
jimczi
left a comment
There was a problem hiding this comment.
Walked through the propagation chain (Lucene → BlobStoreCacheDirectory.openInput → BlobCacheIndexInput.slice(4-arg) → CacheFileReader.copyWithContext → AdvisingRangeMissingHandler.fillCacheRange → SharedBytes.IO.madvise) and the hint plumbing looks right. Role-gating + feature flag mean indexing nodes and prod clusters stay on MADV_NORMAL, and the compound-file boundary math (roundUp/roundDown, adviceForRange falling back to MADV_NORMAL on straddling reads) checks out. Two things worth sorting before you benchmark — the first is probably the story behind whatever numbers you get.
1. Advice only gets set on cache misses, so warm regions never see MADV_RANDOM
AdvisingRangeMissingHandler.fillCacheRange is the only place that calls channel.madvise(advice), and populateAndRead only invokes the RangeMissingHandler when a region is missing. The prewarm/warming paths (SearchCommitPrefetcher, SharedBlobCacheWarmingService, StatelessOnlinePrewarmingService) all go through SequentialRangeMissingHandler / LazyRangeMissingHandler — none of them wrap with AdvisingRangeMissingHandler. So on a search node:
prewarm fills the
.cfsregion (advice stays atMADV_NORMAL) → first KNN query hits the warm cache → reader callback runs,RangeMissingHandlerdoes not →madvise(MADV_RANDOM)is never issued
Cold-cache benches will look great, warm-cache (the steady state) won't see anything. You already have the primitive to fix this — SharedBlobCacheService.adviseRegions(offset, length, advice) walks resident regions and SharedBytes.IO.madvise is idempotent. Calling it once from copyWithContext over the exclusive range would cover already-resident regions; new fills keep flowing through AdvisingRangeMissingHandler. Syscall cost stays at once-per-slice-open.
2. testVecFileGetsRandomAccessHint is broken
It does dir.openInput("_0.vec", IOContext.DEFAULT) and asserts the captured context contains DataAccessHint.RANDOM, but BlobStoreCacheDirectory.openInput just passes the context through — there's no filename-based injection on this side, the vector readers in server/ are what add the hint. IOContext.DEFAULT.hints() is empty, so this is assertTrue(false). testNonVecFileDoesNotGetHint and testExplicitHintIsPreserved pass but for trivial reasons. I'd just drop the three — the propagation they're trying to verify lives in Lucene and is covered there.
3. Region advice can leak across tenants (low risk, same fix as #1)
SharedBytes.IO.currentAdvice is per physical region and regions get reused at eviction. If region R was last advised MADV_RANDOM for shard A's .vec and gets re-tenanted by shard B's .doc via the warming path (which doesn't touch AdvisingRangeMissingHandler), the kernel keeps the stale advice until something writes new advice. Same adviseRegions-from-copyWithContext fix covers it.
|
@jimczi Thanks for the detailed response:
I'll iterate on this and post a note when done. |
|
Thanks for the thorough review, @jimczi . All three points addressed:
I also relaxed the volatile guard in SharedBytes.IO.madvise to opaque VarHandle access, to avoid a memory fence on every read while still skipping redundant syscalls (a stale read can only cause a harmless duplicate madvise). I also added additional unit tests to cover the prewarm scenarios. |
| public void madvise(int advice) { | ||
| if (mmap && (int) VH_CURRENT_ADVICE.getOpaque(this) != advice) { | ||
| mappedByteBuffer.madvise(0, regionSize, advice); | ||
| VH_CURRENT_ADVICE.setOpaque(this, advice); | ||
| } | ||
| } |
There was a problem hiding this comment.
To confirm my understanding, we never expect to call mappedByteBuffer.madvise(0, regionSize, advice); with different values of advice. Further, we don't call mappedByteBuffer.madvise(0, regionSize, NORMAL) because that is the default in this class and in the OS. We only call for RANDOM and in the event of a race, we might call multiple times on the same region (which is no problem as explained in the description and in the comment above).
michaeljmarshall
left a comment
There was a problem hiding this comment.
Looks good to me, but I'm still too new to this part of the code to give a formal approval. I left some minor comments.
| // Actual advice may differ per read — see adviceForRange(). | ||
| private final int desiredMAdvice; | ||
|
|
||
| // The byte range [exclusiveStart, exclusiveEnd) within the blob where cache regions |
There was a problem hiding this comment.
Nit: I found the usage of exclusive here a bit confusing because I assumed they referred to the boundary handling.
| // Passes the sub-file's absolute offset and length so that CacheFileReader can | ||
| // compute which cache regions are exclusively this sub-file's data (safe for | ||
| // MADV_RANDOM) versus boundary regions shared with adjacent sub-files (MADV_NORMAL). | ||
| long subFileOffset = this.offset + offset; |
There was a problem hiding this comment.
Checking understanding, this works because we're slicing within a slice over the original cacheFileReader, right?
…lastic#150066) I've been investigating whether we can reduce unnecessary kernel read-ahead on blob cache regions that back vector data files (.vec). During KNN search, Lucene reads individual vectors by ordinal — an access pattern that is effectively random. By default, the kernel assumes sequential access (MADV_NORMAL) and speculatively reads ahead, wasting I/O bandwidth and polluting the page cache with data that won't be used. This PR adds madvise(MADV_RANDOM) support to the blob cache read path. When Lucene opens a .vec file, it passes an IOContext with DataAccessHint.RANDOM. I thread that hint through BlobStoreCacheDirectory and BlobCacheIndexInput into CacheFileReader, which translates it into a MADV_RANDOM advice on the underlying mmap'd cache regions. The tricky part is compound files. Multiple logical files (e.g. .vec, .vex, .vem) are packed into a single .cfs blob and share the same cache regions. We can't blindly apply MADV_RANDOM to an entire region if it also contains graph data that benefits from sequential read-ahead. So we compute an exclusive range — the interior region-aligned byte range where a sub-file is the sole occupant: exclusiveStart = the first region boundary at or after the sub-file's start offset exclusiveEnd = the last region boundary at or before the sub-file's end offset Reads that fall entirely within [exclusiveStart, exclusiveEnd) get MADV_RANDOM Reads that touch boundary regions (shared with adjacent sub-files) fall back to MADV_NORMAL This is computed once in CacheFileReader.copyWithContext() when Lucene opens a compound sub-file slice, so there's no per-read allocation or computation beyond a simple range check. I gated MADV_RANDOM on the search role — contextToAdvice only returns MADV_RANDOM when the node has hasSearchRole. On index nodes, even if a vector writer opens a .vec file with DataAccessHint.RANDOM (e.g. during merge), we preserve sequential read-ahead by falling back to MADV_NORMAL The actual madvise syscall is issued via AdvisingRangeMissingHandler, a decorator that wraps the existing RangeMissingHandler and calls channel.madvise(advice) just before populating each cache region. The SharedBytes.IO.madvise() method tracks the current advice per region and skips the syscall if it hasn't changed. I've also added MadviseAdvice constants and a madvise(offset, length, advice) method to the native access layer (CloseableMappedByteBuffer), with tests. The whole thing is behind a feature flag (blob_cache_madvise_random) — enabled on snapshot builds for benchmarking, disabled in production. Override with -Des.blob_cache_madvise_random_feature_flag_enabled=true.
I've been investigating whether we can reduce unnecessary kernel read-ahead on blob cache regions that back vector data files (.vec). During KNN search, Lucene reads individual vectors by ordinal — an access pattern that is effectively random. By default, the kernel assumes sequential access (MADV_NORMAL) and speculatively reads ahead, wasting I/O bandwidth and polluting the page cache with data that won't be used.
This PR adds
madvise(MADV_RANDOM)support to the blob cache read path. When Lucene opens a .vec file, it passes anIOContextwithDataAccessHint.RANDOM. I thread that hint throughBlobStoreCacheDirectoryandBlobCacheIndexInputintoCacheFileReader, which translates it into aMADV_RANDOMadvice on the underlying mmap'd cache regions.The tricky part is compound files. Multiple logical files (e.g. .vec, .vex, .vem) are packed into a single .cfs blob and share the same cache regions. We can't blindly apply
MADV_RANDOMto an entire region if it also contains graph data that benefits from sequential read-ahead. So we compute an exclusive range — the interior region-aligned byte range where a sub-file is the sole occupant:This is computed once in
CacheFileReader.copyWithContext()when Lucene opens a compound sub-file slice, so there's no per-read allocation or computation beyond a simple range check.I gated
MADV_RANDOMon the search role —contextToAdviceonly returnsMADV_RANDOMwhen the node hashasSearchRole. On index nodes, even if a vector writer opens a.vecfile withDataAccessHint.RANDOM(e.g. during merge), we preserve sequential read-ahead by falling back toMADV_NORMALThe actual madvise syscall is issued via
AdvisingRangeMissingHandler, a decorator that wraps the existingRangeMissingHandlerand callschannel.madvise(advice)just before populating each cache region. TheSharedBytes.IO.madvise()method tracks the current advice per region and skips the syscall if it hasn't changed.I've also added
MadviseAdviceconstants and amadvise(offset, length, advice)method to the native access layer (CloseableMappedByteBuffer), with tests.The whole thing is behind a feature flag (
blob_cache_madvise_random) — enabled on snapshot builds for benchmarking, disabled in production. Override with-Des.blob_cache_madvise_random_feature_flag_enabled=true.Benchmarking to follow.
relates #147625