Repository navigation
Fix CCR follow to handle indexing_complete race - #145304
Conversation
Fix and unmute. The current fix is adding extra intermediate assertBusy to effectively wait for longer if settings propagation takes logner than 30s. We do similarly in other test cases. Relates elastic#132050 and elastic#137565
|
Pinging @elastic/es-distributed (Team:Distributed) |
DiannaHohensee
left a comment
There was a problem hiding this comment.
It doesn't seem like anyone knows what the problem is yet? I don't think we should increase time when the problem isn't understood. Adding more asserts, and more logging, would be a good thing, though.
Exclude index.lifecycle.indexing_complete from the validateSettings() comparison. This is safe because the shard follow task's innerUpdateSettings() uses filter() independently and will replicate the setting to the follower once following starts. When ILM sets index.lifecycle.indexing_complete on a leader index during rollover, the auto-follow restore may capture stale settings without that flag. The subsequent ResumeFollowAction.validateSettings() then rejects the follow request because the leader and follower settings differ, permanently preventing the shard follow task from being created. Since the error is caught by a fire-and-forget listener with debug-level logging, the failure is silent and unrecoverable. Closes elastic#132050
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 3 out of 3 changed files in this pull request and generated no new comments.
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| .put(filter(leaderIndexSettings)) | ||
| // index.lifecycle.indexing_complete can legitimately differ at follow startup: the leader's ILM may | ||
| // set it during the restore window, before the shard follow task starts syncing settings. | ||
| .remove(LifecycleSettings.LIFECYCLE_INDEXING_COMPLETE) |
There was a problem hiding this comment.
It looks to me that the setting should be added to NON_REPLICATED_SETTINGS instead?
There was a problem hiding this comment.
Adding it to NON_REPLICATED_SETTINGS would stop it from replicating from leader to follower (code). But we want to replicate it. It makes the test fail on every run as well.
So we need two slightly different filters for PutFollow and ResumeFollow. But it could be abstracted into a separate filter method if need.
There was a problem hiding this comment.
After checking the code it seems that ILM settings are indeed replicated between leader to follower, which I find surprising given the follower might have its own ILM policy.
But it seems that in that case, the ILM policy on the follower should have an "unfollow" or a "wait for indexing complete" step before some other destructive actions. So kind of OK if policies are correctly set up.
tlrx
left a comment
There was a problem hiding this comment.
LGTM, though I wonder if this is a legit >bug rather than a >test? IIUC it can happen on production clusters.
| .put(filter(leaderIndexSettings)) | ||
| // index.lifecycle.indexing_complete can legitimately differ at follow startup: the leader's ILM may | ||
| // set it during the restore window, before the shard follow task starts syncing settings. | ||
| .remove(LifecycleSettings.LIFECYCLE_INDEXING_COMPLETE) |
There was a problem hiding this comment.
After checking the code it seems that ILM settings are indeed replicated between leader to follower, which I find surprising given the follower might have its own ILM policy.
But it seems that in that case, the ILM policy on the follower should have an "unfollow" or a "wait for indexing complete" step before some other destructive actions. So kind of OK if policies are correctly set up.
|
Removing myself from review, no need for mine with Tanguy's approval. |
|
Hi @yungene, I've created a changelog YAML for you. |
🔍 Preview links for changed docs⏳ Building and deploying preview... View progress This comment will be updated with preview links when the build is complete. |
ℹ️ Important: Docs version tagging👋 Thanks for updating the docs! Just a friendly reminder that our docs are now cumulative. This means all 9.x versions are documented on the same page and published off of the main branch, instead of creating separate pages for each minor version. We use applies_to tags to mark version-specific features and changes. Expand for a quick overviewWhen to use applies_to tags:✅ At the page level to indicate which products/deployments the content applies to (mandatory) What NOT to do:❌ Don't remove or replace information that applies to an older version 🤔 Need help?
|
This reverts commit 58eb00e.
Exclude index.lifecycle.indexing_complete from the validateSettings() comparison. This is safe because the shard follow task's innerUpdateSettings() uses filter() independently and will replicate the setting to the follower once following starts. When ILM sets index.lifecycle.indexing_complete on a leader index during rollover, the auto-follow restore may capture stale settings without that flag. The subsequent ResumeFollowAction.validateSettings() then rejects the follow request because the leader and follower settings differ, permanently preventing the shard follow task from being created. Since the error is caught by a fire-and-forget listener with debug-level logging, the failure is silent and unrecoverable. Relates #132050 and #137565
Exclude index.lifecycle.indexing_complete from the validateSettings() comparison. This is safe because the shard follow task's innerUpdateSettings() uses filter() independently and will replicate the setting to the follower once following starts. When ILM sets index.lifecycle.indexing_complete on a leader index during rollover, the auto-follow restore may capture stale settings without that flag. The subsequent ResumeFollowAction.validateSettings() then rejects the follow request because the leader and follower settings differ, permanently preventing the shard follow task from being created. Since the error is caught by a fire-and-forget listener with debug-level logging, the failure is silent and unrecoverable. Relates elastic#132050 and elastic#137565
* Fix CCR follow to handle indexing_complete race (#145304) Exclude index.lifecycle.indexing_complete from the validateSettings() comparison. This is safe because the shard follow task's innerUpdateSettings() uses filter() independently and will replicate the setting to the follower once following starts. When ILM sets index.lifecycle.indexing_complete on a leader index during rollover, the auto-follow restore may capture stale settings without that flag. The subsequent ResumeFollowAction.validateSettings() then rejects the follow request because the leader and follower settings differ, permanently preventing the shard follow task from being created. Since the error is caught by a fire-and-forget listener with debug-level logging, the failure is silent and unrecoverable. Relates #132050 and #137565 * Backport settings remove
* Fix CCR follow to handle indexing_complete race (#145304) Exclude index.lifecycle.indexing_complete from the validateSettings() comparison. This is safe because the shard follow task's innerUpdateSettings() uses filter() independently and will replicate the setting to the follower once following starts. When ILM sets index.lifecycle.indexing_complete on a leader index during rollover, the auto-follow restore may capture stale settings without that flag. The subsequent ResumeFollowAction.validateSettings() then rejects the follow request because the leader and follower settings differ, permanently preventing the shard follow task from being created. Since the error is caught by a fire-and-forget listener with debug-level logging, the failure is silent and unrecoverable. Relates #132050 and #137565 * Backport settings remove
Exclude index.lifecycle.indexing_complete from the validateSettings() comparison. This is safe because the shard follow task's innerUpdateSettings() uses filter() independently and will replicate the setting to the follower once following starts. When ILM sets index.lifecycle.indexing_complete on a leader index during rollover, the auto-follow restore may capture stale settings without that flag. The subsequent ResumeFollowAction.validateSettings() then rejects the follow request because the leader and follower settings differ, permanently preventing the shard follow task from being created. Since the error is caught by a fire-and-forget listener with debug-level logging, the failure is silent and unrecoverable. Relates elastic#132050 and elastic#137565
Exclude index.lifecycle.indexing_complete from the validateSettings() comparison. This is safe because the shard follow task's innerUpdateSettings() uses filter() independently and will replicate the setting to the follower once following starts. When ILM sets index.lifecycle.indexing_complete on a leader index during rollover, the auto-follow restore may capture stale settings without that flag. The subsequent ResumeFollowAction.validateSettings() then rejects the follow request because the leader and follower settings differ, permanently preventing the shard follow task from being created. Since the error is caught by a fire-and-forget listener with debug-level logging, the failure is silent and unrecoverable. Relates elastic#132050 and elastic#137565
Exclude index.lifecycle.indexing_complete from the validateSettings()
comparison. This is safe because the shard follow task's
innerUpdateSettings() uses filter() independently and will replicate
the setting to the follower once following starts.
When ILM sets index.lifecycle.indexing_complete on a leader index
during rollover, the auto-follow restore may capture stale settings
without that flag. The subsequent ResumeFollowAction.validateSettings()
then rejects the follow request because the leader and follower
settings differ, permanently preventing the shard follow task from
being created. Since the error is caught by a fire-and-forget listener
with debug-level logging, the failure is silent and unrecoverable.
Relates #132050 and #137565