Repository navigation
Desired-balance warn threshold logging should accumulate across restarts #100850
Copy link
Copy link
Closed
Closed
Copy link
Labels
:Distributed/AllocationAll issues relating to the decision making around placing a shard (both master logic & on the nodes)All issues relating to the decision making around placing a shard (both master logic & on the nodes)>enhancementSupportabilityImprove our (devs, SREs, support eng, users) ability to troubleshoot/self-service product better.Improve our (devs, SREs, support eng, users) ability to troubleshoot/self-service product better.Team:DistributedMeta label for distributed team.Meta label for distributed team.
Description
Activity
- added:Distributed/AllocationAll issues relating to the decision making around placing a shard (both master logic & on the nodes)All issues relating to the decision making around placing a shard (both master logic & on the nodes)SupportabilityImprove our (devs, SREs, support eng, users) ability to troubleshoot/self-service product better.Improve our (devs, SREs, support eng, users) ability to troubleshoot/self-service product better.
on Oct 13, 2023 - addedTeam:DistributedMeta label for distributed team.Meta label for distributed team.
on Oct 13, 2023 elasticsearchmachine commented
on Oct 13, 2023 CollaboratorMore actionsPinging @elastic/es-distributed (Team:Distributed)
Related to #91386
- addedTeam:Distributed Coordination (obsolete)Meta label for Distributed Coordination team. Obsolete. Please do not use.Meta label for Distributed Coordination team. Obsolete. Please do not use.
on Mar 26, 2025 elasticsearchmachine commented
on Mar 26, 2025 CollaboratorMore actionsPinging @elastic/es-distributed-obsolete (Team:Distributed (Obsolete))
elasticsearchmachine commented
on Mar 26, 2025 CollaboratorMore actionsPinging @elastic/es-distributed-coordination (Team:Distributed Coordination)
- added 4 commits that reference this issue
on Apr 1, 2025 - added a commit that references this issue
on Apr 9, 2025 - added a commit that references this issue
on Oct 31, 2025 - added a commit that references this issue
on Dec 5, 2025 - removedTeam:Distributed Coordination (obsolete)Meta label for Distributed Coordination team. Obsolete. Please do not use.Meta label for Distributed Coordination team. Obsolete. Please do not use.
on Feb 18, 2026
Metadata
Metadata
Assignees
Labels
:Distributed/AllocationAll issues relating to the decision making around placing a shard (both master logic & on the nodes)All issues relating to the decision making around placing a shard (both master logic & on the nodes)>enhancementSupportabilityImprove our (devs, SREs, support eng, users) ability to troubleshoot/self-service product better.Improve our (devs, SREs, support eng, users) ability to troubleshoot/self-service product better.Team:DistributedMeta label for distributed team.Meta label for distributed team.
Today we emit periodic INFO logs about an ongoing desired balance computation which has not converged after some amount of time or some number of iterations:
elasticsearch/server/src/main/java/org/elasticsearch/cluster/routing/allocation/allocator/DesiredBalanceComputer.java
Lines 321 to 329 in 18f960c
However these numbers reset if a new cluster state is received, so it's possible for a steady stream of cluster states to prevent the computation for converging without ever seeing any warnings. IMO we should let these numbers accumulate until the computation fully converges, and report the number of restarts since convergence in the log message too.