Skip to content

Advise MADV_RANDOM on blob cache regions backing vector data files - #150066

Merged
ChrisHegarty merged 39 commits into
elastic:mainfrom
ChrisHegarty:madvise-random-vec-blob-cache
Jun 18, 2026
Merged

ChrisHegarty merged 39 commits into
elastic:mainfrom
ChrisHegarty:madvise-random-vec-blob-cache

Conversation

@ChrisHegarty

@ChrisHegarty ChrisHegarty commented May 28, 2026 •

Copy link
Copy Markdown
Contributor

I've been investigating whether we can reduce unnecessary kernel read-ahead on blob cache regions that back vector data files (.vec). During KNN search, Lucene reads individual vectors by ordinal — an access pattern that is effectively random. By default, the kernel assumes sequential access (MADV_NORMAL) and speculatively reads ahead, wasting I/O bandwidth and polluting the page cache with data that won't be used.

This PR adds madvise(MADV_RANDOM) support to the blob cache read path. When Lucene opens a .vec file, it passes an IOContext with DataAccessHint.RANDOM. I thread that hint through BlobStoreCacheDirectory and BlobCacheIndexInput into CacheFileReader, which translates it into a MADV_RANDOM advice on the underlying mmap'd cache regions.

The tricky part is compound files. Multiple logical files (e.g. .vec, .vex, .vem) are packed into a single .cfs blob and share the same cache regions. We can't blindly apply MADV_RANDOM to an entire region if it also contains graph data that benefits from sequential read-ahead. So we compute an exclusive range — the interior region-aligned byte range where a sub-file is the sole occupant:

  • exclusiveStart = the first region boundary at or after the sub-file's start offset
  • exclusiveEnd = the last region boundary at or before the sub-file's end offset
  • Reads that fall entirely within [exclusiveStart, exclusiveEnd) get MADV_RANDOM
  • Reads that touch boundary regions (shared with adjacent sub-files) fall back to MADV_NORMAL

This is computed once in CacheFileReader.copyWithContext() when Lucene opens a compound sub-file slice, so there's no per-read allocation or computation beyond a simple range check.

I gated MADV_RANDOM on the search role — contextToAdvice only returns MADV_RANDOM when the node has hasSearchRole. On index nodes, even if a vector writer opens a .vec file with DataAccessHint.RANDOM (e.g. during merge), we preserve sequential read-ahead by falling back to MADV_NORMAL

The actual madvise syscall is issued via AdvisingRangeMissingHandler, a decorator that wraps the existing RangeMissingHandler and calls channel.madvise(advice) just before populating each cache region. The SharedBytes.IO.madvise() method tracks the current advice per region and skips the syscall if it hasn't changed.

I've also added MadviseAdvice constants and a madvise(offset, length, advice) method to the native access layer (CloseableMappedByteBuffer), with tests.

The whole thing is behind a feature flag (blob_cache_madvise_random) — enabled on snapshot builds for benchmarking, disabled in production. Override with -Des.blob_cache_madvise_random_feature_flag_enabled=true.

Benchmarking to follow.

relates #147625

ChrisHegarty and others added 16 commits April 27, 2026 14:23
Remove the separate StatelessKnnIndexTester class and the x-pack/qa/vector
subproject. The stateless directory type is now registered directly in
KnnIndexTester static initializer, alongside default and frozen.

The stateless directory factory and cache stats logger are invoked via
reflection (same pattern as the existing frozen/searchable-snapshots
directory) since the stateless test artifact lives in an unnamed module.
…ix graph stats

- Add skip_recall config option to skip brute-force exact NN
  computation for QPS-only benchmark runs
- Fix race condition in KnnIndexer where ID_FIELD could get out of
  sync with the actual vector read from the dataset file by moving
  ordinal assignment into IndexVectorReader
- Fix logGraphStats to unwrap PerFieldKnnVectorsFormat.FieldsReader
  before checking for HnswGraphProvider

Co-authored-by: Cursor <cursoragent@cursor.com>
@ChrisHegarty ChrisHegarty added >feature :Search Relevance/Vectors Vector search Team:Search Relevance Meta label for the Search Relevance team in Elasticsearch labels May 28, 2026
@elasticsearchmachine

Copy link
Copy Markdown
Collaborator

Pinging @elastic/es-search-relevance (Team:Search Relevance)

@elasticsearchmachine

Copy link
Copy Markdown
Collaborator

Hi @ChrisHegarty, I've created a changelog YAML for you.

@github-actions

github-actions Bot commented May 28, 2026 •

Copy link
Copy Markdown
Contributor

🔍 Preview links for changed docs

⏳ Building and deploying preview... View progress

This comment will be updated with preview links when the build is complete.

@github-actions

Copy link
Copy Markdown
Contributor

ℹ️ Important: Docs version tagging

👋 Thanks for updating the docs! Just a friendly reminder that our docs are now cumulative. This means all 9.x versions are documented on the same page and published off of the main branch, instead of creating separate pages for each minor version.

We use applies_to tags to mark version-specific features and changes.

Expand for a quick overview

When to use applies_to tags:

✅ At the page level to indicate which products/deployments the content applies to (mandatory)
✅ When features change state (e.g. preview, ga) in a specific version
✅ When availability differs across deployments and environments

What NOT to do:

❌ Don't remove or replace information that applies to an older version
❌ Don't add new information that applies to a specific version without an applies_to tag
❌ Don't forget that applies_to tags can be used at the page, section, and inline level

🤔 Need help?

@ChrisHegarty

Copy link
Copy Markdown
Contributor Author

For non-compound (top-level) files, each file exclusively owns its blob — there are no shared regions. The exclusive range is [0, Long.MAX_VALUE), so every region gets MADV_RANDOM. No boundary fallback is needed.

For compound files (.cfs), multiple sub-files are packed into a single blob and share cache regions at the boundaries. The table below shows how significant (or insignificant) those boundary regions are relative to the total .vec file size, using the default blob cache region size of 16 MiB:

Dims Encoding Bytes/vec Vectors .vec size Total regions Boundary regions Boundary %
384 byte 384 100K 36.6 MiB 3 2 66.7%
384 byte 384 1M 366 MiB 23 2 8.7%
384 byte 384 10M 3.6 GiB 229 2 0.9%
768 byte 768 1M 732 MiB 46 2 4.3%
768 byte 768 10M 7.2 GiB 457 2 0.4%
1024 byte 1024 1M 976 MiB 61 2 3.3%
1024 byte 1024 10M 9.5 GiB 610 2 0.3%
1024 float32 4096 1M 3.8 GiB 244 2 0.8%
1024 float32 4096 10M 38.1 GiB 2441 2 0.08%

At most 2 boundary regions (32 MiB) per sub-file fall back to MADV_NORMAL — one at each end where the sub-file shares a cache region with an adjacent file in the compound blob. For any non-trivial dataset (1M+ vectors), this is well under 5% of regions. For typical production workloads (millions of vectors, 768+ dims), over 99% of regions receive MADV_RANDOM. The small dataset case (100K vectors, 384 dims) is the worst case at 66.7% boundary — but at only 36.6 MiB total, the entire file fits comfortably in cache and read-ahead suppression matters far less.

@jimczi jimczi left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Walked through the propagation chain (Lucene → BlobStoreCacheDirectory.openInput → BlobCacheIndexInput.slice(4-arg) → CacheFileReader.copyWithContext → AdvisingRangeMissingHandler.fillCacheRange → SharedBytes.IO.madvise) and the hint plumbing looks right. Role-gating + feature flag mean indexing nodes and prod clusters stay on MADV_NORMAL, and the compound-file boundary math (roundUp/roundDown, adviceForRange falling back to MADV_NORMAL on straddling reads) checks out. Two things worth sorting before you benchmark — the first is probably the story behind whatever numbers you get.

1. Advice only gets set on cache misses, so warm regions never see MADV_RANDOM

AdvisingRangeMissingHandler.fillCacheRange is the only place that calls channel.madvise(advice), and populateAndRead only invokes the RangeMissingHandler when a region is missing. The prewarm/warming paths (SearchCommitPrefetcher, SharedBlobCacheWarmingService, StatelessOnlinePrewarmingService) all go through SequentialRangeMissingHandler / LazyRangeMissingHandler — none of them wrap with AdvisingRangeMissingHandler. So on a search node:

prewarm fills the .cfs region (advice stays at MADV_NORMAL) → first KNN query hits the warm cache → reader callback runs, RangeMissingHandler does not → madvise(MADV_RANDOM) is never issued

Cold-cache benches will look great, warm-cache (the steady state) won't see anything. You already have the primitive to fix this — SharedBlobCacheService.adviseRegions(offset, length, advice) walks resident regions and SharedBytes.IO.madvise is idempotent. Calling it once from copyWithContext over the exclusive range would cover already-resident regions; new fills keep flowing through AdvisingRangeMissingHandler. Syscall cost stays at once-per-slice-open.

2. testVecFileGetsRandomAccessHint is broken

It does dir.openInput("_0.vec", IOContext.DEFAULT) and asserts the captured context contains DataAccessHint.RANDOM, but BlobStoreCacheDirectory.openInput just passes the context through — there's no filename-based injection on this side, the vector readers in server/ are what add the hint. IOContext.DEFAULT.hints() is empty, so this is assertTrue(false). testNonVecFileDoesNotGetHint and testExplicitHintIsPreserved pass but for trivial reasons. I'd just drop the three — the propagation they're trying to verify lives in Lucene and is covered there.

3. Region advice can leak across tenants (low risk, same fix as #1)

SharedBytes.IO.currentAdvice is per physical region and regions get reused at eviction. If region R was last advised MADV_RANDOM for shard A's .vec and gets re-tenanted by shard B's .doc via the warming path (which doesn't touch AdvisingRangeMissingHandler), the kernel keeps the stale advice until something writes new advice. Same adviseRegions-from-copyWithContext fix covers it.

@ChrisHegarty

Copy link
Copy Markdown
Contributor Author

@jimczi Thanks for the detailed response:

  1. I was not aware of the hot versus cold distinction at this level. I had assumed that data gets into the cache in the same way for both, so the advise would work as is for both. Let me recheck this and fix it.
  2. Yeah, testVecFileGetsRandomAccessHint, is was a bad unit test. I since removed it.
  3. Same as no1. - bad assumption.

I'll iterate on this and post a note when done.

@ChrisHegarty

ChrisHegarty commented May 29, 2026 •

Copy link
Copy Markdown
Contributor Author

Thanks for the thorough review, @jimczi . All three points addressed:

  1. Warm regions. I moved madvise out of AdvisingRangeMissingHandler (fill path only) and into the reader callback on both populateAndRead and tryRead. Advice is now applied on every read regardless of how the region was populated. AdvisingRangeMissingHandler is removed.
  2. Broken test. Already removed.
  3. Stale advice on reuse — same fix as no.1; the read path overwrites whatever the previous tenant left behind.

I also relaxed the volatile guard in SharedBytes.IO.madvise to opaque VarHandle access, to avoid a memory fence on every read while still skipping redundant syscalls (a stale read can only cause a harmless duplicate madvise).

I also added additional unit tests to cover the prewarm scenarios.

@ChrisHegarty
ChrisHegarty requested a review from jimczi May 29, 2026 10:00
Comment on lines +392 to +397
public void madvise(int advice) {
if (mmap && (int) VH_CURRENT_ADVICE.getOpaque(this) != advice) {
mappedByteBuffer.madvise(0, regionSize, advice);
VH_CURRENT_ADVICE.setOpaque(this, advice);
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To confirm my understanding, we never expect to call mappedByteBuffer.madvise(0, regionSize, advice); with different values of advice. Further, we don't call mappedByteBuffer.madvise(0, regionSize, NORMAL) because that is the default in this class and in the OS. We only call for RANDOM and in the event of a race, we might call multiple times on the same region (which is no problem as explained in the description and in the comment above).

@michaeljmarshall michaeljmarshall left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good to me, but I'm still too new to this part of the code to give a formal approval. I left some minor comments.

// Actual advice may differ per read — see adviceForRange().
private final int desiredMAdvice;

// The byte range [exclusiveStart, exclusiveEnd) within the blob where cache regions

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: I found the usage of exclusive here a bit confusing because I assumed they referred to the boundary handling.

Comment on lines +124 to +127
// Passes the sub-file's absolute offset and length so that CacheFileReader can
// compute which cache regions are exclusively this sub-file's data (safe for
// MADV_RANDOM) versus boundary regions shared with adjacent sub-files (MADV_NORMAL).
long subFileOffset = this.offset + offset;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Checking understanding, this works because we're slicing within a slice over the original cacheFileReader, right?

@jimczi jimczi left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@ChrisHegarty
ChrisHegarty merged commit a72d563 into elastic:main Jun 18, 2026
37 checks passed
Kubik42 pushed a commit to Kubik42/elasticsearch that referenced this pull request Jun 18, 2026
…lastic#150066)

I've been investigating whether we can reduce unnecessary kernel read-ahead on blob cache regions that back vector data files (.vec). During KNN search, Lucene reads individual vectors by ordinal — an access pattern that is effectively random. By default, the kernel assumes sequential access (MADV_NORMAL) and speculatively reads ahead, wasting I/O bandwidth and polluting the page cache with data that won't be used.

This PR adds madvise(MADV_RANDOM) support to the blob cache read path. When Lucene opens a .vec file, it passes an IOContext with DataAccessHint.RANDOM. I thread that hint through BlobStoreCacheDirectory and BlobCacheIndexInput into CacheFileReader, which translates it into a MADV_RANDOM advice on the underlying mmap'd cache regions.

The tricky part is compound files. Multiple logical files (e.g. .vec, .vex, .vem) are packed into a single .cfs blob and share the same cache regions. We can't blindly apply MADV_RANDOM to an entire region if it also contains graph data that benefits from sequential read-ahead. So we compute an exclusive range — the interior region-aligned byte range where a sub-file is the sole occupant:

exclusiveStart = the first region boundary at or after the sub-file's start offset
exclusiveEnd = the last region boundary at or before the sub-file's end offset
Reads that fall entirely within [exclusiveStart, exclusiveEnd) get MADV_RANDOM
Reads that touch boundary regions (shared with adjacent sub-files) fall back to MADV_NORMAL
This is computed once in CacheFileReader.copyWithContext() when Lucene opens a compound sub-file slice, so there's no per-read allocation or computation beyond a simple range check.

I gated MADV_RANDOM on the search role — contextToAdvice only returns MADV_RANDOM when the node has hasSearchRole. On index nodes, even if a vector writer opens a .vec file with DataAccessHint.RANDOM (e.g. during merge), we preserve sequential read-ahead by falling back to MADV_NORMAL

The actual madvise syscall is issued via AdvisingRangeMissingHandler, a decorator that wraps the existing RangeMissingHandler and calls channel.madvise(advice) just before populating each cache region. The SharedBytes.IO.madvise() method tracks the current advice per region and skips the syscall if it hasn't changed.

I've also added MadviseAdvice constants and a madvise(offset, length, advice) method to the native access layer (CloseableMappedByteBuffer), with tests.

The whole thing is behind a feature flag (blob_cache_madvise_random) — enabled on snapshot builds for benchmarking, disabled in production. Override with -Des.blob_cache_madvise_random_feature_flag_enabled=true.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

>feature :Search Relevance/Vectors Vector search Team:Search Relevance Meta label for the Search Relevance team in Elasticsearch v9.5.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants