Repository navigation
GPU codec: fall back to CPU graph build on flush when GPU is busy - #149373
Conversation
When flushing, use tryAcquire (non-blocking) to attempt GPU resource acquisition. If the GPU is busy, fall back to building the HNSW graph on CPU using HnswGraphBuilder. This avoids blocking flush threads waiting for GPU resources during heavy indexing. Also adds a `reason` parameter to acquire/tryAcquire for improved diagnostics, and refactors both methods to share a common doAcquire implementation.
|
Pinging @elastic/es-search-relevance (Team:Search Relevance) |
|
Hi @ChrisHegarty, I've created a changelog YAML for you. |
🔍 Preview links for changed docs⏳ Building and deploying preview... View progress This comment will be updated with preview links when the build is complete. |
ℹ️ Important: Docs version tagging👋 Thanks for updating the docs! Just a friendly reminder that our docs are now cumulative. This means all 9.x versions are documented on the same page and published off of the main branch, instead of creating separate pages for each minor version. We use applies_to tags to mark version-specific features and changes. Expand for a quick overviewWhen to use applies_to tags:✅ At the page level to indicate which products/deployments the content applies to (mandatory) What NOT to do:❌ Don't remove or replace information that applies to an older version 🤔 Need help?
|
| ); | ||
| logger.error(message); | ||
| throw new IllegalArgumentException(message); | ||
| throw memoryExceededError(numVectors, dims, totalMemoryInBytes); |
There was a problem hiding this comment.
can we also fallback to CPU for this case: when there is not enough memory on GPU?
There was a problem hiding this comment.
we can also leave it for a follow up
There was a problem hiding this comment.
This should never be true now, right? If the memory exceeds what the GPU has, we switch algorithm.
I'm ok with the change as a safety net, in case we have data so large and a GPU so small even IVF-PQ cannot handle them, but it should be a follow-up
There was a problem hiding this comment.
Right. It's an effective runtime assert. ( I do not want to make it an assert as I want it to always be surfaced, since it's an error )
mayya-sharipova
left a comment
There was a problem hiding this comment.
Thanks Chris for fixing this issue
ldematte
left a comment
There was a problem hiding this comment.
I like the idea. I think it's very smart to have this in flush, were data should be on the smaller side, but keep acquire on merge.
| ); | ||
| logger.error(message); | ||
| throw new IllegalArgumentException(message); | ||
| throw memoryExceededError(numVectors, dims, totalMemoryInBytes); |
There was a problem hiding this comment.
This should never be true now, right? If the memory exceeds what the GPU has, we switch algorithm.
I'm ok with the change as a safety net, in case we have data so large and a GPU so small even IVF-PQ cannot handle them, but it should be a follow-up
| if (numLockedResources() == 0) { | ||
| logger.debug("No resources currently locked, proceeding"); | ||
| // If no resource in the pool is locked, we must proceed to avoid livelock | ||
| if (enoughMemory == false && numLockedResources() == 0) { |
There was a problem hiding this comment.
Why has the enoughMemory check been moved here?
There was a problem hiding this comment.
I added the enoughMemory == false check here so that when a resource is available and there is enough memory, then the regular call flow is followed - not this short-circuit. The regular acquire call flow, when enoughMemory is true, will later log the acquisition and update the memory service.
…49373) (#149475) During heavy indexing, multiple flush operations can compete for GPU resources simultaneously. Previously, flush would block waiting for a GPU resource to become available, which stalls the indexing thread and can cause cascading latency. This is particularly problematic when the GPU is already saturated with merge or other flush operations — the thread just sits idle waiting for its turn. I've changed the flush path to use a non-blocking tryAcquire instead of a blocking acquire. If the GPU is busy (all resources locked or insufficient memory), flush now builds the HNSW graph on CPU using Lucene's HnswGraphBuilder. The resulting graph is written in the same Lucene99 format, so it's fully searchable by the standard reader. This means flush never blocks on GPU availability — it always makes progress, just potentially slower for that particular relatively small new segment. To support this, I added tryAcquire to the CuVSResourceManager interface. Both acquire and tryAcquire now delegate to a shared doAcquire implementation with a nonBlocking flag, and I've added a reason parameter for diagnostics so we can see in logs which operation is acquiring or waiting for resources. I've added tests at three levels: unit tests for the tryAcquire mechanics (including a concurrent contention test), a WriteGraphTests class that validates the CPU fallback produces byte-identical output to Lucene, and two mixed-path format tests that exercise both GPU and CPU paths within the same index on GPU nodes via a randomly-failing resource manager.
…sy (#149373) (#149476) * GPU codec: fall back to CPU graph build on flush when GPU is busy (#149373) During heavy indexing, multiple flush operations can compete for GPU resources simultaneously. Previously, flush would block waiting for a GPU resource to become available, which stalls the indexing thread and can cause cascading latency. This is particularly problematic when the GPU is already saturated with merge or other flush operations — the thread just sits idle waiting for its turn. I've changed the flush path to use a non-blocking tryAcquire instead of a blocking acquire. If the GPU is busy (all resources locked or insufficient memory), flush now builds the HNSW graph on CPU using Lucene's HnswGraphBuilder. The resulting graph is written in the same Lucene99 format, so it's fully searchable by the standard reader. This means flush never blocks on GPU availability — it always makes progress, just potentially slower for that particular relatively small new segment. To support this, I added tryAcquire to the CuVSResourceManager interface. Both acquire and tryAcquire now delegate to a shared doAcquire implementation with a nonBlocking flag, and I've added a reason parameter for diagnostics so we can see in logs which operation is acquiring or waiting for resources. I've added tests at three levels: unit tests for the tryAcquire mechanics (including a concurrent contention test), a WriteGraphTests class that validates the CPU fallback produces byte-identical output to Lucene, and two mixed-path format tests that exercise both GPU and CPU paths within the same index on GPU nodes via a randomly-failing resource manager. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix compilation for Lucene 10.3.2 Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com>
During heavy indexing, multiple flush operations can compete for GPU resources simultaneously. Previously, flush would block waiting for a GPU resource to become available, which stalls the indexing thread and can cause cascading latency. This is particularly problematic when the GPU is already saturated with merge or other flush operations — the thread just sits idle waiting for its turn.
I've changed the flush path to use a non-blocking
tryAcquireinstead of a blockingacquire. If the GPU is busy (all resources locked or insufficient memory),flushnow builds the HNSW graph on CPU using Lucene'sHnswGraphBuilder. The resulting graph is written in the same Lucene99 format, so it's fully searchable by the standard reader. This meansflushnever blocks on GPU availability — it always makes progress, just potentially slower for that particular relatively small new segment.To support this, I added
tryAcquireto theCuVSResourceManagerinterface. BothacquireandtryAcquirenow delegate to a shareddoAcquireimplementation with a nonBlocking flag, and I've added areasonparameter for diagnostics so we can see in logs which operation is acquiring or waiting for resources.I've added tests at three levels: unit tests for the
tryAcquiremechanics (including a concurrent contention test), aWriteGraphTestsclass that validates the CPU fallback produces byte-identical output to Lucene, and two mixed-path format tests that exercise both GPU and CPU paths within the same index on GPU nodes via a randomly-failing resource manager.