Repository navigation
[beats receivers] Connection level output 401 error handling does not match the process runtime #14531
Description
Activity
In the Elasticsearch exporter, the same
RetryOnStatusconfiguration applies to both request and document level retries making it so that we can't exactly match the beats behavior which would retry indefinitely for completely invalid credentials but not stall the pipeline because of a single missing index privilege in an otherwise valid API key (which would be a document level 401):- addedTeam:Elastic-Agent-Control-PlaneLabel for the Agent Control Plane teamLabel for the Agent Control Plane teamTeam:Elastic-Agent-Data-PlaneLabel for the Agent Data Plane teamLabel for the Agent Data Plane team
on May 25, 2026 infra-vault-gh-plugin-prod commented
on May 25, 2026 More actionsPinging @elastic/elastic-agent-control-plane (Team:Elastic-Agent-Control-Plane)
infra-vault-gh-plugin-prod commented
on May 25, 2026 More actionsPinging @elastic/elastic-agent-data-plane (Team:Elastic-Agent-Data-Plane)
Short term it is probably better to retry on request and document level 401+403 if we can fix this quickly.
It does look like the ES exporter would allow setting separate request and document level retry configurations:
return docappender.BulkIndexerConfig{ Client: client, MaxDocumentRetries: maxDocRetries, Pipeline: config.Pipeline, RetryOnDocumentStatus: config.Retry.RetryOnStatus,
RetryOnStatus: config.Retry.RetryOnStatus, DisableRetry: !config.Retry.Enabled, RetryOnError: func(_ *http.Request, err error) bool { return !errors.Is(err, context.Canceled) && !errors.Is(err, context.DeadlineExceeded) },
github-actions commented
on May 25, 2026 on May 25, 2026 – with GitHub ActionsContributorMore actionstl;dr I can reproduce the mismatch in code: the OTel Elasticsearch output translation used by beats-receiver/OTel runtime explicitly excludes
401from retryable status codes, while the process runtime path still uses libbeat output handling that retries connection-level failures.Recommendation
Align beats-receiver/OTel runtime behavior with process runtime for connection-level401by updating the OTel ES translation retry policy and adding a regression test. Concretely, updateinternal/pkg/otel/translate/output_elasticsearch.goso retry behavior for401matches the agreed policy for process runtime parity, and lock it with tests inoutput_elasticsearch_test.go+otelconfig_test.go.Findings
-
OTel translation currently hard-codes retry statuses to
429+5xxonly:internal/pkg/otel/translate/output_elasticsearch.go:38-53(RetryOnStatusdefault list; no401)internal/pkg/otel/translate/output_elasticsearch.go:180-188(writesretry.retry_on_statusfrom that list)
-
Tests assert and preserve this behavior today:
internal/pkg/otel/translate/output_elasticsearch_test.go:59-71internal/pkg/otel/translate/otelconfig_test.go:291-297
-
Runtime paths are split, so this translation only affects OTel-managed components:
internal/pkg/agent/application/coordinator/coordinator.go:2035-2042(runtime model + otel model split/update)internal/pkg/agent/application/coordinator/coordinator.go:2199-2207(splitModelBetweenManagersby runtime manager)
-
Process runtime output unit still passes output config directly (no OTel translation layer):
pkg/component/component.go:605-612(ExpectedConfig(output.Config)for output unit)
-
In libbeat Elasticsearch client, connection-level errors are retried via
handleBulkResultError, while document-level<500statuses are generally non-retry (except 429):beats/libbeat/outputs/elasticsearch/client.go:258-261,327-359beats/libbeat/outputs/elasticsearch/client.go:519-526
-
Related upstream context:
elastic/beats#50261(merged): “fix(otelconsumer): do not retry 401 errors from Elasticsearch”
Verification
I validated current behavior by running targeted translation tests.
$ cd /home/runner/work/elastic-agent/elastic-agent $ go test ./internal/pkg/otel/translate -run 'TestToOtelConfig|TestCompressionConfig|TestToOTelConfig_CheckUnsupported' go: go.mod requires go >= 1.26.3 (running go 1.25.10; GOTOOLCHAIN=local)
$ cd /home/runner/work/elastic-agent/elastic-agent $ GOTOOLCHAIN=auto go test ./internal/pkg/otel/translate -run 'TestToOtelConfig|TestCompressionConfig|TestToOTelConfig_CheckUnsupported' ok github.com/elastic/elastic-agent/internal/pkg/otel/translate 0.230s
$ cd /home/runner/work/elastic-agent/elastic-agent $ GOTOOLCHAIN=auto go test ./internal/pkg/otel/translate -run TestUnitToExporterConfig ok github.com/elastic/elastic-agent/internal/pkg/otel/translate 0.312s
These passing tests currently encode the default retry list without
401.Detailed Action Plan
-
Update OTel ES translation retry defaults in
internal/pkg/otel/translate/output_elasticsearch.go:- Modify
defaultOptions.RetryOnStatus(aroundL38-L53) to match the intended parity policy for connection-level401.
- Modify
-
Update translation expectations in tests:
internal/pkg/otel/translate/output_elasticsearch_test.goexpected YAMLretry_on_statuslists.internal/pkg/otel/translate/otelconfig_test.go:291-297expectedretry_on_statuslist.
-
Add a focused regression test that asserts parity intent explicitly:
- In
output_elasticsearch_test.go, add a test case that validates presence/absence of401in translated retry policy per the decided behavior.
- In
-
Validate with targeted package tests:
GOTOOLCHAIN=auto go test ./internal/pkg/otel/translate -run 'TestToOtelConfig|TestUnitToExporterConfig'
-
Optional hardening (if team wants stronger guardrail):
- Add a coordinator/runtime-level test that ensures OTel-managed output retry policy remains aligned with process-runtime policy for
401handling assumptions.
- Add a coordinator/runtime-level test that ensures OTel-managed output retry policy remains aligned with process-runtime policy for
Related Items
Type Link / File Relevance Issue #14531 Current triage target PR elastic/beats#50261 Upstream change touching 401 retry behavior in otelconsumer File internal/pkg/otel/translate/output_elasticsearch.go:38-53Default retryable statuses (no 401) File internal/pkg/otel/translate/output_elasticsearch.go:180-188Maps retry config into OTel exporter File internal/pkg/otel/translate/output_elasticsearch_test.go:59-71Test fixtures asserting current retry list File internal/pkg/otel/translate/otelconfig_test.go:291-297Expected retry list in unit-to-exporter mapping File pkg/component/component.go:605-612Process runtime output unit config path File internal/pkg/agent/application/coordinator/coordinator.go:2035-2042Runtime/OTel split update path File beats/libbeat/outputs/elasticsearch/client.go:258-261Connection-level error handling path File beats/libbeat/outputs/elasticsearch/client.go:519-526Per-item non-retry behavior for <500except 429Note
🔒 Integrity filter blocked 3 items
The following items were blocked because they don't meet the GitHub integrity level.
- #401
search_pull_requests: has lower integrity than agent requires. The agent cannot read data with integrity below "approved". - #50217
issue_read: has lower integrity than agent requires. The agent cannot read data with integrity below "approved". - #9406
issue_read: has lower integrity than agent requires. The agent cannot read data with integrity below "approved".
To allow these resources, lower
min-integrityin your GitHub frontmatter:tools: github: min-integrity: approved # merged | approved | unapproved | none
What is this? | From workflow: Issue Triage
Give us feedback! React with 🚀 if perfect, 👍 if helpful, 👎 if not.
-
Linking to the upstream feature request open-telemetry/opentelemetry-collector-contrib#48681
- linked a pull request that will close this issue[elasticsearchexporter] Add retry_on_document_status configuration #48933
on Jun 9, 2026 When we implement this on the agent side, let's make sure we have a test that the exponential backoff on 401 works with the expected delay increase. There is an active support case right now where versions before 9.3.5 retry immediately causing force unenrolled agents to hammer Elasticsearch and we want to avoid introducing that bug again.
@cmacknz should we maintain the current status code retried in Otel mode and add 401 and 403 for request level retries, or should we fully match the Elasticsearch output and retry anything >= 500 at document level? Currently I'm going with the first approach, but that's an easy change.
We should exactly match the beats implementation, so if I've specified anything that conflicts with that go with the beats implementation (unless there is a flaw in it, then we should fix both variants).
Reacted by Tiago Queiroz- added a commit that references this issue
on Aug 21, 2026 - added a commit that references this issue
on Aug 21, 2026 - added 2 commits that reference this issue
on Aug 24, 2026
In the beats Elasticsearch output, connection level errors like 401s are always retried:
https://github.com/elastic/beats/blob/71d0f09b540201d44155615d01a35639895d64f3/libbeat/outputs/elasticsearch/client.go#L255-L264
However, most document level 4xx codes like 401s are not retried:
https://github.com/elastic/beats/blob/71d0f09b540201d44155615d01a35639895d64f3/libbeat/outputs/elasticsearch/client.go#L491-L523
In Elastic Agent, the ES output translation configures the following set of retryable status codes which excludes 401s:
elastic-agent/internal/pkg/otel/translate/output_elasticsearch.go
Lines 38 to 53 in 902c65b
This difference is causing problems in other systems like ECK where TLS certificate propagation has some delay, leading to data loss on the first events shipped after startup. See elastic/cloud-on-k8s#9406