Skip to content

Fix CCR follow to handle indexing_complete race - #145304

Merged
yungene merged 21 commits into
elastic:mainfrom
yungene:2026/03/31/ccr-test-fix
May 7, 2026
Merged

yungene merged 21 commits into
elastic:mainfrom
yungene:2026/03/31/ccr-test-fix

Conversation

@yungene

@yungene yungene commented Mar 31, 2026 •

Copy link
Copy Markdown
Contributor

Exclude index.lifecycle.indexing_complete from the validateSettings()
comparison. This is safe because the shard follow task's
innerUpdateSettings() uses filter() independently and will replicate
the setting to the follower once following starts.

When ILM sets index.lifecycle.indexing_complete on a leader index
during rollover, the auto-follow restore may capture stale settings
without that flag. The subsequent ResumeFollowAction.validateSettings()
then rejects the follow request because the leader and follower
settings differ, permanently preventing the shard follow task from
being created. Since the error is caught by a fire-and-forget listener
with debug-level logging, the failure is silent and unrecoverable.

Relates #132050 and #137565

Fix and unmute. The current fix is adding extra intermediate assertBusy
to effectively wait for longer if settings propagation takes logner than
30s. We do similarly in other test cases.

Relates elastic#132050 and
elastic#137565

This comment was marked as outdated.

@yungene
yungene marked this pull request as ready for review April 2, 2026 09:11
@elasticsearchmachine elasticsearchmachine added the needs:triage Requires assignment of a team area label label Apr 2, 2026
@yungene yungene added >test Issues or PRs that are addressing/adding tests :Distributed/CCR Issues around the Cross Cluster State Replication features and removed needs:triage Requires assignment of a team area label labels Apr 2, 2026
@yungene
yungene requested a review from DiannaHohensee April 2, 2026 09:12
@elasticsearchmachine elasticsearchmachine added the Team:Distributed Meta label for distributed team. label Apr 2, 2026
@elasticsearchmachine

Copy link
Copy Markdown
Collaborator

Pinging @elastic/es-distributed (Team:Distributed)

@DiannaHohensee DiannaHohensee left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It doesn't seem like anyone knows what the problem is yet? I don't think we should increase time when the problem isn't understood. Adding more asserts, and more logging, would be a good thing, though.

yungene added 2 commits April 14, 2026 15:39
Exclude index.lifecycle.indexing_complete from the validateSettings()
comparison. This is safe because the shard follow task's
innerUpdateSettings() uses filter() independently and will replicate
the setting to the follower once following starts.

When ILM sets index.lifecycle.indexing_complete on a leader index
during rollover, the auto-follow restore may capture stale settings
without that flag. The subsequent ResumeFollowAction.validateSettings()
then rejects the follow request because the leader and follower
settings differ, permanently preventing the shard follow task from
being created. Since the error is caught by a fire-and-forget listener
with debug-level logging, the failure is silent and unrecoverable.

Closes elastic#132050
@yungene yungene changed the title Unmute testTsdbLeaderIndexRolloverAndSyncAfterWaitUntilEndTime Fix CCR follow to handle indexing_complete race Apr 14, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 3 out of 3 changed files in this pull request and generated no new comments.


💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

.put(filter(leaderIndexSettings))
// index.lifecycle.indexing_complete can legitimately differ at follow startup: the leader's ILM may
// set it during the restore window, before the shard follow task starts syncing settings.
.remove(LifecycleSettings.LIFECYCLE_INDEXING_COMPLETE)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It looks to me that the setting should be added to NON_REPLICATED_SETTINGS instead?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Adding it to NON_REPLICATED_SETTINGS would stop it from replicating from leader to follower (code). But we want to replicate it. It makes the test fail on every run as well.

So we need two slightly different filters for PutFollow and ResumeFollow. But it could be abstracted into a separate filter method if need.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

After checking the code it seems that ILM settings are indeed replicated between leader to follower, which I find surprising given the follower might have its own ILM policy.

But it seems that in that case, the ILM policy on the follower should have an "unfollow" or a "wait for indexing complete" step before some other destructive actions. So kind of OK if policies are correctly set up.

@tlrx tlrx left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, though I wonder if this is a legit >bug rather than a >test? IIUC it can happen on production clusters.

.put(filter(leaderIndexSettings))
// index.lifecycle.indexing_complete can legitimately differ at follow startup: the leader's ILM may
// set it during the restore window, before the shard follow task starts syncing settings.
.remove(LifecycleSettings.LIFECYCLE_INDEXING_COMPLETE)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

After checking the code it seems that ILM settings are indeed replicated between leader to follower, which I find surprising given the follower might have its own ILM policy.

But it seems that in that case, the ILM policy on the follower should have an "unfollow" or a "wait for indexing complete" step before some other destructive actions. So kind of OK if policies are correctly set up.

@DiannaHohensee

Copy link
Copy Markdown
Contributor

Removing myself from review, no need for mine with Tanguy's approval.

@DiannaHohensee
DiannaHohensee removed their request for review April 29, 2026 23:00
@yungene yungene added >bug and removed >test Issues or PRs that are addressing/adding tests labels May 5, 2026
@elasticsearchmachine

Copy link
Copy Markdown
Collaborator

Hi @yungene, I've created a changelog YAML for you.

@github-actions

github-actions Bot commented May 5, 2026 •

Copy link
Copy Markdown
Contributor

🔍 Preview links for changed docs

⏳ Building and deploying preview... View progress

This comment will be updated with preview links when the build is complete.

@github-actions

github-actions Bot commented May 5, 2026

Copy link
Copy Markdown
Contributor

ℹ️ Important: Docs version tagging

👋 Thanks for updating the docs! Just a friendly reminder that our docs are now cumulative. This means all 9.x versions are documented on the same page and published off of the main branch, instead of creating separate pages for each minor version.

We use applies_to tags to mark version-specific features and changes.

Expand for a quick overview

When to use applies_to tags:

✅ At the page level to indicate which products/deployments the content applies to (mandatory)
✅ When features change state (e.g. preview, ga) in a specific version
✅ When availability differs across deployments and environments

What NOT to do:

❌ Don't remove or replace information that applies to an older version
❌ Don't add new information that applies to a specific version without an applies_to tag
❌ Don't forget that applies_to tags can be used at the page, section, and inline level

🤔 Need help?

@yungene
yungene merged commit 55813e3 into elastic:main May 7, 2026
39 checks passed
elasticsearchmachine pushed a commit that referenced this pull request May 8, 2026
Exclude index.lifecycle.indexing_complete from the validateSettings()
comparison. This is safe because the shard follow task's
innerUpdateSettings() uses filter() independently and will replicate
the setting to the follower once following starts.

When ILM sets index.lifecycle.indexing_complete on a leader index
during rollover, the auto-follow restore may capture stale settings
without that flag. The subsequent ResumeFollowAction.validateSettings()
then rejects the follow request because the leader and follower
settings differ, permanently preventing the shard follow task from
being created. Since the error is caught by a fire-and-forget listener
with debug-level logging, the failure is silent and unrecoverable.

Relates #132050 and #137565
ncordon pushed a commit to ncordon/elasticsearch that referenced this pull request May 8, 2026
Exclude index.lifecycle.indexing_complete from the validateSettings()
comparison. This is safe because the shard follow task's
innerUpdateSettings() uses filter() independently and will replicate
the setting to the follower once following starts.

When ILM sets index.lifecycle.indexing_complete on a leader index
during rollover, the auto-follow restore may capture stale settings
without that flag. The subsequent ResumeFollowAction.validateSettings()
then rejects the follow request because the leader and follower
settings differ, permanently preventing the shard follow task from
being created. Since the error is caught by a fire-and-forget listener
with debug-level logging, the failure is silent and unrecoverable.

Relates elastic#132050 and elastic#137565
elasticsearchmachine pushed a commit that referenced this pull request May 8, 2026
* Fix CCR follow to handle indexing_complete race (#145304)

Exclude index.lifecycle.indexing_complete from the validateSettings()
comparison. This is safe because the shard follow task's
innerUpdateSettings() uses filter() independently and will replicate
the setting to the follower once following starts.

When ILM sets index.lifecycle.indexing_complete on a leader index
during rollover, the auto-follow restore may capture stale settings
without that flag. The subsequent ResumeFollowAction.validateSettings()
then rejects the follow request because the leader and follower
settings differ, permanently preventing the shard follow task from
being created. Since the error is caught by a fire-and-forget listener
with debug-level logging, the failure is silent and unrecoverable.

Relates #132050 and #137565

* Backport settings remove
elasticsearchmachine pushed a commit that referenced this pull request May 8, 2026
* Fix CCR follow to handle indexing_complete race (#145304)

Exclude index.lifecycle.indexing_complete from the validateSettings()
comparison. This is safe because the shard follow task's
innerUpdateSettings() uses filter() independently and will replicate
the setting to the follower once following starts.

When ILM sets index.lifecycle.indexing_complete on a leader index
during rollover, the auto-follow restore may capture stale settings
without that flag. The subsequent ResumeFollowAction.validateSettings()
then rejects the follow request because the leader and follower
settings differ, permanently preventing the shard follow task from
being created. Since the error is caught by a fire-and-forget listener
with debug-level logging, the failure is silent and unrecoverable.

Relates #132050 and #137565

* Backport settings remove
alighahramani-alig pushed a commit to alighahramani-alig/elasticsearch that referenced this pull request May 8, 2026
Exclude index.lifecycle.indexing_complete from the validateSettings()
comparison. This is safe because the shard follow task's
innerUpdateSettings() uses filter() independently and will replicate
the setting to the follower once following starts.

When ILM sets index.lifecycle.indexing_complete on a leader index
during rollover, the auto-follow restore may capture stale settings
without that flag. The subsequent ResumeFollowAction.validateSettings()
then rejects the follow request because the leader and follower
settings differ, permanently preventing the shard follow task from
being created. Since the error is caught by a fire-and-forget listener
with debug-level logging, the failure is silent and unrecoverable.

Relates elastic#132050 and elastic#137565
henningandersen pushed a commit to henningandersen/elasticsearch that referenced this pull request May 11, 2026
Exclude index.lifecycle.indexing_complete from the validateSettings()
comparison. This is safe because the shard follow task's
innerUpdateSettings() uses filter() independently and will replicate
the setting to the follower once following starts.

When ILM sets index.lifecycle.indexing_complete on a leader index
during rollover, the auto-follow restore may capture stale settings
without that flag. The subsequent ResumeFollowAction.validateSettings()
then rejects the follow request because the leader and follower
settings differ, permanently preventing the shard follow task from
being created. Since the error is caught by a fire-and-forget listener
with debug-level logging, the failure is silent and unrecoverable.

Relates elastic#132050 and elastic#137565
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

>bug :Distributed/CCR Issues around the Cross Cluster State Replication features Team:Distributed Meta label for distributed team. v9.5.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants