Skip to content

[Transform] Honor cluster health wait timeout when ensuring system index shards are active - #149462

Merged
prwhelan merged 10 commits into
elastic:mainfrom
rnb-tron:fix/149400-cluster-health-timeout
Jun 15, 2026
Merged

prwhelan merged 10 commits into
elastic:mainfrom
rnb-tron:fix/149400-cluster-health-timeout

Conversation

@rnb-tron

Copy link
Copy Markdown
Contributor

Summary

Fix TransformInternalIndex#waitForLatestVersionedIndexShardsActive to
treat a ClusterHealth timeout as a failure instead of silently
completing the listener with success.

Closes #149400

Background

When a transform is created via PUT _transform/{id}, the action calls
TransformInternalIndex#createLatestVersionedIndexIfRequired. If the
versioned internal index already exists but its shards are not yet
active, the code falls back to waitForLatestVersionedIndexShardsActive,
which issues a ClusterHealth request that waits up to 30 seconds for
at least one active shard.

The previous implementation handled the response like this:

ActionListener<ClusterHealthResponse> innerListener = ActionListener.wrap(
    r -> listener.onResponse(null),   // always success!
    listener::onFailure
);

Any non-exceptional response — including timeouts where
ClusterHealthResponse#isTimedOut() returns true — was forwarded
straight to onResponse(null). As a result the caller assumed the
system index was ready and proceeded to write the transform config,
which would later surface as confusing downstream failures.

What changed

In TransformInternalIndex#waitForLatestVersionedIndexShardsActive:

  • Inspect response.isTimedOut() before completing the listener.
  • On timeout, fail the listener with
    ElasticsearchStatusException(SERVICE_UNAVAILABLE) carrying a
    dedicated message constant so the REST layer returns an actionable
    503 and the client can retry.
  • Refactor to the project's idiomatic listener.delegateFailureAndWrap
    style (consistent with IndexBasedTransformConfigManager and many
    other call sites in the module).

In TransformMessages:

  • Add a focused constant
    TRANSFORM_WAIT_FOR_INDEX_SHARDS_ACTIVE_TIMEOUT = "Timed out waiting for transform system index shards to become active".

Tests

Added TransformInternalIndexTests#testCreateLatestVersionedIndexIfRequired_GivenClusterHealthTimeout
which:

  1. Sets up a cluster state where the versioned index exists but shards
    are not active, forcing the code into the waitFor… branch.
  2. Stubs the cluster admin client to return a ClusterHealthResponse
    with setTimedOut(true).
  3. Asserts that the listener is failed with
    ElasticsearchStatusException, that
    status() == RestStatus.SERVICE_UNAVAILABLE, and that the message
    matches TransformMessages.TRANSFORM_WAIT_FOR_INDEX_SHARDS_ACTIVE_TIMEOUT.
  4. Verifies the cluster health client is invoked exactly once.
./gradlew :x-pack:plugin:transform:test \
  --tests "org.elasticsearch.xpack.transform.persistence.TransformInternalIndexTests"

Behavior change

Scenario Before After
Shards become active within 30s Success Success (unchanged)
Inner request fails onFailure propagated onFailure propagated (unchanged)
Cluster health request times out Silently onResponse(null) onFailure(503 SERVICE_UNAVAILABLE)

This is a user-visible behavior change: callers that previously
appeared to succeed (and then failed later in a confusing way) will now
receive a clear 503 they can retry. No transport version or
serialization changes are required.

`waitForLatestVersionedIndexShardsActive` issues a `ClusterHealth`
request that waits up to 30s for at least one active shard of the
transform internal index. The listener previously dispatched any
non-exceptional response straight to `onResponse(null)`, including
responses where `isTimedOut() == true`. As a result a PUT transform
call could proceed to write its config even when the system index
shards were not actually ready, surfacing later as obscure failures.

Treat `isTimedOut() == true` as a failure and complete the listener
with an `ElasticsearchStatusException(SERVICE_UNAVAILABLE)` so the
client gets an actionable 503 and can retry.

Closes elastic#149400
@elasticsearchmachine elasticsearchmachine added needs:triage Requires assignment of a team area label v9.5.0 external-contributor Pull request authored by a developer outside the Elasticsearch team labels May 20, 2026
@kingherc kingherc added the :ml/Transform Transform label May 20, 2026
@elasticsearchmachine elasticsearchmachine added Team:ML Meta label for the ML team and removed needs:triage Requires assignment of a team area label labels May 20, 2026
@elasticsearchmachine

Copy link
Copy Markdown
Collaborator

Pinging @elastic/ml-core (Team:ML)

@prwhelan prwhelan left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good - thank you for the contribution

@prwhelan
prwhelan enabled auto-merge (squash) June 4, 2026 19:15
@prwhelan

prwhelan commented Jun 5, 2026

Copy link
Copy Markdown
Member

In queue, will merge Monday June 15th

@elasticsearchmachine

Copy link
Copy Markdown
Collaborator

💚 Backport successful

Status Branch Result
✅ 8.19
✅ 9.3
✅ 9.4

elasticsearchmachine pushed a commit that referenced this pull request Jun 15, 2026
…dex shards are active (#149462) (#151226)

`waitForLatestVersionedIndexShardsActive` issues a `ClusterHealth`
request that waits up to 30s for at least one active shard of the
transform internal index. The listener previously dispatched any
non-exceptional response straight to `onResponse(null)`, including
responses where `isTimedOut() == true`. As a result a PUT transform
call could proceed to write its config even when the system index
shards were not actually ready, surfacing later as obscure failures.

Treat `isTimedOut() == true` as a failure and complete the listener
with an `ElasticsearchStatusException(SERVICE_UNAVAILABLE)` so the
client gets an actionable 503 and can retry.

Closes #149400

Co-authored-by: Tron Wang <59386838+rnb-tron@users.noreply.github.com>
elasticsearchmachine pushed a commit that referenced this pull request Jun 15, 2026
…dex shards are active (#149462) (#151228)

`waitForLatestVersionedIndexShardsActive` issues a `ClusterHealth`
request that waits up to 30s for at least one active shard of the
transform internal index. The listener previously dispatched any
non-exceptional response straight to `onResponse(null)`, including
responses where `isTimedOut() == true`. As a result a PUT transform
call could proceed to write its config even when the system index
shards were not actually ready, surfacing later as obscure failures.

Treat `isTimedOut() == true` as a failure and complete the listener
with an `ElasticsearchStatusException(SERVICE_UNAVAILABLE)` so the
client gets an actionable 503 and can retry.

Closes #149400

Co-authored-by: Tron Wang <59386838+rnb-tron@users.noreply.github.com>
elasticsearchmachine pushed a commit that referenced this pull request Jun 15, 2026
…dex shards are active (#149462) (#151227)

`waitForLatestVersionedIndexShardsActive` issues a `ClusterHealth`
request that waits up to 30s for at least one active shard of the
transform internal index. The listener previously dispatched any
non-exceptional response straight to `onResponse(null)`, including
responses where `isTimedOut() == true`. As a result a PUT transform
call could proceed to write its config even when the system index
shards were not actually ready, surfacing later as obscure failures.

Treat `isTimedOut() == true` as a failure and complete the listener
with an `ElasticsearchStatusException(SERVICE_UNAVAILABLE)` so the
client gets an actionable 503 and can retry.

Closes #149400

Co-authored-by: Tron Wang <59386838+rnb-tron@users.noreply.github.com>
valeriy42 pushed a commit to valeriy42/elasticsearch that referenced this pull request Jun 18, 2026
…dex shards are active (elastic#149462)

`waitForLatestVersionedIndexShardsActive` issues a `ClusterHealth`
request that waits up to 30s for at least one active shard of the
transform internal index. The listener previously dispatched any
non-exceptional response straight to `onResponse(null)`, including
responses where `isTimedOut() == true`. As a result a PUT transform
call could proceed to write its config even when the system index
shards were not actually ready, surfacing later as obscure failures.

Treat `isTimedOut() == true` as a failure and complete the listener
with an `ElasticsearchStatusException(SERVICE_UNAVAILABLE)` so the
client gets an actionable 503 and can retry.

Closes elastic#149400

Co-authored-by: Pat Whelan <pat.whelan@elastic.co>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

auto-backport Automatically create backport pull requests when merged >bug external-contributor Pull request authored by a developer outside the Elasticsearch team :ml/Transform Transform Team:ML Meta label for the ML team v8.19.17 v9.3.6 v9.4.3 v9.5.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Transform] waitForLatestVersionedIndexShardsActive ignores ClusterHealth timeout

4 participants