Skip to content

Improve bulker performance - #7190

Merged
michel-laterman merged 5 commits into
elastic:mainfrom
michel-laterman:worktree-bulk-benchmarks
Jun 17, 2026
Merged

michel-laterman merged 5 commits into
elastic:mainfrom
michel-laterman:worktree-bulk-benchmarks

Conversation

@michel-laterman

Copy link
Copy Markdown
Contributor

What is the problem this PR solves?

Improve bulker performance by lowering the number of allocations.

How does this PR solve the problem?

APM Improvements:

APM span handling has changed from using a pointer to a directly allocated span per bulkT (removing 1 heap allocation)

Replaced withAPMLinkedContext with direct inline extraction (avoids allocating a slice and closure)

Pre-allocate span links

Pre-compute span name arrays

Removed span in Read operation -> read always calls ReadRaw

Bug fix

Correctly trace calls in opMulti.go

Other improvements

OpApiKey.go - use queue.cnt as a size hint to pre-allocate space

flushUpdateAPIKey - replace json.Decoder call that discards decoded bytes with a bytes.Cut in order to discard metadata.

How to test this PR locally

Additional benchmarks have been added with dae7d64.

A benchstat comparison between dae7d64 and 11ff6f2 produced the following results:

pkg: github.com/elastic/fleet-server/v7/internal/pkg/bulk
                           │ build/benchmark-baseline.out │       build/benchmark-new.out        │
                           │            sec/op            │    sec/op     vs base                │
MockBulk/1-12                                39.87µ ±  3%   37.81µ ±  1%   -5.17% (p=0.000 n=10)
MockBulk/8-12                                140.0µ ±  1%   134.8µ ±  1%   -3.68% (p=0.000 n=10)
MockBulk/64-12                               779.9µ ±  4%   768.7µ ±  1%        ~ (p=0.075 n=10)
MockBulk/4096-12                             55.96m ±  2%   57.92m ±  2%   +3.49% (p=0.005 n=10)
MockBulk/32768-12                            516.4m ±  4%   517.0m ±  1%        ~ (p=0.971 n=10)
DispatchAbortQueue-12                        108.9n ±  1%   107.0n ±  2%   -1.70% (p=0.016 n=10)
DispatchAbortResponse-12                     368.5n ± 18%   365.0n ± 20%        ~ (p=0.912 n=10)
DispatchSuccess-12                           413.1n ±  2%   410.2n ±  2%        ~ (p=0.190 n=10)
FlushSearch/1-12                             2.365µ ±  6%   1.696µ ±  5%  -28.29% (p=0.000 n=10)
FlushSearch/8-12                             5.776µ ±  1%   5.096µ ±  1%  -11.77% (p=0.000 n=10)
FlushSearch/64-12                            34.79µ ±  4%   30.62µ ±  1%  -11.99% (p=0.000 n=10)
FlushSearch/4096-12                          2.181m ±  2%   1.895m ±  5%  -13.08% (p=0.000 n=10)
FlushSearch/32768-12                         17.66m ±  3%   15.27m ±  1%  -13.51% (p=0.000 n=10)
FlushRead/1-12                               2.068µ ±  6%   1.418µ ±  2%  -31.41% (p=0.000 n=10)
FlushRead/8-12                               3.836µ ±  6%   3.070µ ±  8%  -19.97% (p=0.000 n=10)
FlushRead/64-12                              19.64µ ±  2%   15.30µ ±  1%  -22.09% (p=0.000 n=10)
FlushRead/4096-12                           1128.8µ ±  3%   884.7µ ± 10%  -21.62% (p=0.000 n=10)
FlushRead/32768-12                           8.912m ±  1%   7.076m ±  2%  -20.61% (p=0.000 n=10)
FlushAPIKeyUpdate/1-12                       3.610µ ±  3%   2.739µ ±  3%  -24.12% (p=0.000 n=10)
FlushAPIKeyUpdate/8-12                      17.710µ ±  1%   9.652µ ±  3%  -45.50% (p=0.000 n=10)
FlushAPIKeyUpdate/64-12                     140.98µ ±  1%   71.39µ ±  2%  -49.36% (p=0.000 n=10)
FlushAPIKeyUpdate/4096-12                    9.059m ±  2%   4.455m ±  1%  -50.82% (p=0.000 n=10)
FlushAPIKeyUpdate/32768-12                   78.15m ± 10%   40.12m ±  3%  -48.67% (p=0.000 n=10)
MultiUpdateMock/1-12                         797.6n ±  2%   751.8n ±  3%   -5.74% (p=0.000 n=10)
MultiUpdateMock/8-12                         3.390µ ±  1%   3.449µ ±  2%   +1.73% (p=0.005 n=10)
MultiUpdateMock/64-12                        18.09µ ±  1%   17.83µ ±  1%   -1.42% (p=0.001 n=10)
MultiUpdateMock/4096-12                      941.6µ ±  2%   926.7µ ±  1%   -1.58% (p=0.003 n=10)
MultiUpdateMock/32768-12                     7.441m ±  2%   7.293m ±  5%        ~ (p=0.353 n=10)
MultiUpdateMock/131072-12                    31.06m ±  1%   31.44m ±  2%   +1.24% (p=0.000 n=10)
geomean                                      123.0µ         102.4µ        -16.75%
                           │ build/benchmark-baseline.out │         build/benchmark-new.out         │
                           │             B/op             │     B/op       vs base                  │
MockBulk/1-12                            36.51Ki ±   0%     29.51Ki ±  0%  -19.16% (p=0.000 n=10)
MockBulk/8-12                            96.07Ki ±   0%     72.62Ki ±  0%  -24.41% (p=0.000 n=10)
MockBulk/64-12                           486.5Ki ±   0%     432.9Ki ±  0%  -11.00% (p=0.000 n=10)
MockBulk/4096-12                         28.89Mi ±   0%     28.28Mi ±  0%   -2.12% (p=0.000 n=10)
MockBulk/32768-12                        234.9Mi ±   0%     231.4Mi ±  2%   -1.49% (p=0.001 n=10)
DispatchAbortQueue-12                      0.000 ±   0%       0.000 ±  0%        ~ (p=1.000 n=10) ¹
DispatchAbortResponse-12                   102.0 ± 118%       114.0 ± 84%        ~ (p=0.725 n=10)
DispatchSuccess-12                         16.00 ±   0%       16.00 ±  0%        ~ (p=1.000 n=10) ¹
FlushSearch/1-12                         4.416Ki ±   0%     2.689Ki ±  0%  -39.10% (p=0.000 n=10)
FlushSearch/8-12                         5.789Ki ±   0%     3.977Ki ±  0%  -31.31% (p=0.000 n=10)
FlushSearch/64-12                        27.90Ki ±   0%     13.40Ki ±  0%  -51.97% (p=0.000 n=10)
FlushSearch/4096-12                     1610.5Ki ±   0%     683.1Ki ±  0%  -57.59% (p=0.000 n=10)
FlushSearch/32768-12                    12.565Mi ±   0%     5.349Mi ±  0%  -57.43% (p=0.000 n=10)
FlushRead/1-12                           4.453Ki ±   0%     2.492Ki ±  0%  -44.04% (p=0.000 n=10)
FlushRead/8-12                           5.227Ki ±   0%     3.148Ki ±  0%  -39.76% (p=0.000 n=10)
FlushRead/64-12                         25.664Ki ±   0%     9.148Ki ±  0%  -64.35% (p=0.000 n=10)
FlushRead/4096-12                       1322.4Ki ±   0%     386.7Ki ±  0%  -70.76% (p=0.000 n=10)
FlushRead/32768-12                      10.260Mi ±   0%     3.018Mi ±  0%  -70.58% (p=0.000 n=10)
FlushAPIKeyUpdate/1-12                   4.523Ki ±   0%     3.031Ki ±  0%  -32.99% (p=0.000 n=10)
FlushAPIKeyUpdate/8-12                  18.336Ki ±   0%     6.398Ki ±  0%  -65.10% (p=0.000 n=10)
FlushAPIKeyUpdate/64-12                 151.30Ki ±   0%     57.96Ki ±  0%  -61.69% (p=0.000 n=10)
FlushAPIKeyUpdate/4096-12                9.518Mi ±   0%     3.692Mi ±  0%  -61.22% (p=0.000 n=10)
FlushAPIKeyUpdate/32768-12               76.96Mi ±   0%     30.25Mi ±  0%  -60.70% (p=0.000 n=10)
MultiUpdateMock/1-12                       472.0 ±   0%       480.0 ±  0%   +1.69% (p=0.000 n=10)
MultiUpdateMock/8-12                     1.977Ki ±   0%     2.219Ki ±  0%  +12.25% (p=0.000 n=10)
MultiUpdateMock/64-12                    14.73Ki ±   0%     15.97Ki ±  0%   +8.44% (p=0.000 n=10)
MultiUpdateMock/4096-12                  880.2Ki ±   0%     976.2Ki ±  0%  +10.91% (p=0.000 n=10)
MultiUpdateMock/32768-12                 6.875Mi ±   0%     7.625Mi ±  0%  +10.91% (p=0.000 n=10)
MultiUpdateMock/131072-12                27.50Mi ±   0%     30.50Mi ±  0%  +10.91% (p=0.000 n=10)
geomean                                                 ²                  -34.32%                ²
¹ all samples are equal
² summaries must be >0 to compute geomean
                           │ build/benchmark-baseline.out │        build/benchmark-new.out        │
                           │          allocs/op           │  allocs/op   vs base                  │
MockBulk/1-12                                212.0 ± 0%      184.0 ± 0%  -13.21% (p=0.000 n=10)
MockBulk/8-12                                468.0 ± 0%      360.0 ± 0%  -23.08% (p=0.000 n=10)
MockBulk/64-12                              2.499k ± 0%     1.720k ± 0%  -31.17% (p=0.000 n=10)
MockBulk/4096-12                           148.42k ± 0%     99.23k ± 0%  -33.15% (p=0.000 n=10)
MockBulk/32768-12                          1224.3k ± 0%     831.4k ± 0%  -32.09% (p=0.000 n=10)
DispatchAbortQueue-12                        0.000 ± 0%      0.000 ± 0%        ~ (p=1.000 n=10) ¹
DispatchAbortResponse-12                     1.000 ± 0%      1.000 ± 0%        ~ (p=1.000 n=10) ¹
DispatchSuccess-12                           1.000 ± 0%      1.000 ± 0%        ~ (p=1.000 n=10) ¹
FlushSearch/1-12                             21.00 ± 0%      19.00 ± 0%   -9.52% (p=0.000 n=10)
FlushSearch/8-12                             26.00 ± 0%      26.00 ± 0%        ~ (p=1.000 n=10) ¹
FlushSearch/64-12                            82.00 ± 0%      82.00 ± 0%        ~ (p=1.000 n=10) ¹
FlushSearch/4096-12                         4.114k ± 0%     4.114k ± 0%        ~ (p=1.000 n=10) ¹
FlushSearch/32768-12                        32.79k ± 0%     32.79k ± 0%        ~ (p=1.000 n=10) ¹
FlushRead/1-12                               22.00 ± 0%      19.00 ± 0%  -13.64% (p=0.000 n=10)
FlushRead/8-12                               27.00 ± 0%      26.00 ± 0%   -3.70% (p=0.000 n=10)
FlushRead/64-12                              83.00 ± 0%      82.00 ± 0%   -1.20% (p=0.000 n=10)
FlushRead/4096-12                           4.115k ± 0%     4.114k ± 0%   -0.02% (p=0.000 n=10)
FlushRead/32768-12                          32.79k ± 0%     32.79k ± 0%   -0.00% (p=0.000 n=10)
FlushAPIKeyUpdate/1-12                       46.00 ± 0%      30.00 ± 0%  -34.78% (p=0.000 n=10)
FlushAPIKeyUpdate/8-12                       238.0 ± 0%      103.0 ± 0%  -56.72% (p=0.000 n=10)
FlushAPIKeyUpdate/64-12                     1888.0 ± 0%      738.0 ± 0%  -60.91% (p=0.000 n=10)
FlushAPIKeyUpdate/4096-12                  118.94k ± 0%     45.18k ± 0%  -62.01% (p=0.000 n=10)
FlushAPIKeyUpdate/32768-12                  951.1k ± 0%     361.3k ± 0%  -62.02% (p=0.000 n=10)
MultiUpdateMock/1-12                         7.000 ± 0%      6.000 ± 0%  -14.29% (p=0.000 n=10)
MultiUpdateMock/8-12                         7.000 ± 0%      6.000 ± 0%  -14.29% (p=0.000 n=10)
MultiUpdateMock/64-12                        7.000 ± 0%      6.000 ± 0%  -14.29% (p=0.000 n=10)
MultiUpdateMock/4096-12                      7.000 ± 0%      6.000 ± 0%  -14.29% (p=0.000 n=10)
MultiUpdateMock/32768-12                     7.000 ± 0%      6.000 ± 0%  -14.29% (p=0.000 n=10)
MultiUpdateMock/131072-12                    7.000 ± 0%      6.000 ± 0%  -14.29% (p=0.000 n=10)
geomean                                                 ²                -21.25%                ²
¹ all samples are equal
² summaries must be >0 to compute geomean

There are major improvements across the bulker except for the Multi operations (due to the bug fix)

Design Checklist

  • I have ensured my design is stateless and will work when multiple fleet-server instances are behind a load balancer.
  • I have or intend to scale test my changes, ensuring it will work reliably with 100K+ agents connected.
  • I have included fail safe mechanisms to limit the load on fleet-server: rate limiting, circuit breakers, caching, load shedding, etc.

Checklist

  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • I have made corresponding change to the default configuration files
  • I have added tests that prove my fix is effective or that my feature works
  • I have added an entry in ./changelog/fragments using the changelog tool

michel-laterman and others added 2 commits June 10, 2026 13:19
…date

BenchmarkFlushSearch, BenchmarkFlushRead, and BenchmarkFlushAPIKeyUpdate
directly exercise the three flush paths that had no allocation baseline.
Each benchmark uses a fixedTransport (zero per-request parsing overhead)
and a pre-built queue so measurements isolate the flush functions themselves.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Replace *apm.SpanLink pointer with value type + hasSpanLink bool on
bulkT and optionsT, eliminating one heap allocation per bulk op.
Remove withAPMLinkedContext closure pattern; extract APM transaction
context inline at each call site instead. Fix multiWaitBulkOp to
actually set span links on queued items (previously called
withAPMLinkedContext but never propagated the result) and add a
caller-level APM span for multi-ops.

Pool flush buffers via flushBufPool sync.Pool across flushBulk,
flushSearch, and flushRead to avoid per-flush buffer allocation.
Pre-allocate span link slices with queue.cnt capacity. Pre-compute
APM span name strings as package-level arrays to eliminate fmt.Sprintf
on every flush. Size-hint all six maps in flushUpdateAPIKey with
queue.cnt. Replace double-decode of NDJSON meta in flushUpdateAPIKey
with bytes.Cut to skip the meta line, then unmarshal only the body.
Remove the redundant outer APM span from Read (ReadRaw already spans).

Benchstat results (n=10, benchtime=3s):
- FlushRead: -22% time, -64% bytes
- FlushAPIKeyUpdate: -49% time, -62% bytes
- FlushSearch: -14% time, -52% bytes
- MockBulk/1: -5% time, -19% bytes, -13% allocs

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@michel-laterman
michel-laterman requested a review from a team as a code owner June 11, 2026 11:35
@michel-laterman
michel-laterman requested a review from samuelvl June 11, 2026 11:35
@michel-laterman michel-laterman added enhancement New feature or request backport-skip Skip notification from the automated backport with mergify Team:Elastic-Agent-Control-Plane Label for the Agent Control Plane team labels Jun 11, 2026
@github-actions

This comment has been minimized.

@michel-laterman

Copy link
Copy Markdown
Contributor Author

I re-ran bechmarks with -memprofile, here are some basic findings:

  • withAPMLinkedContext removed from heap (2.75GB in base tests) - this is an obvious change as that method was removed
  • encoding/json.NewDecoder removed (1.75GB in base) - this was replaced by the bytes.Cut
  • multi* operations have more heap usage due to being correctly instrumented now

Overall the removal of withAPMLinkedContext and less JSON decoders should result in less GC intervention.

@github-actions

This comment has been minimized.

michel-laterman and others added 2 commits June 16, 2026 12:52
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@michel-laterman
michel-laterman force-pushed the worktree-bulk-benchmarks branch from a1e4af6 to 738a2f4 Compare June 16, 2026 21:09
@github-actions

This comment has been minimized.

@github-actions

Copy link
Copy Markdown
Contributor

TL;DR

E2E Test failed with go test exit status 1, but the provided Buildkite artifact only contains the final 149 lines of output. In that tail, every listed TestStandAloneRunningSuite subtest passed and there is no assertion, panic, timeout, or race-detector message to trace to this PR.

Remediation

  • Retry the E2E job or fetch the full Buildkite log for fleet-server-e2e-test; the current artifact is missing the actual failing test/error.
  • If the rerun fails, preserve the full go test -v -race -timeout 30m ./... output so the failing suite and stack trace are visible.
Investigation details

Root Cause

Classification: Inconclusive / missing failure data.

The local artifact at /tmp/gh-aw/buildkite-logs/fleet-server-e2e-test.txt is only 149 lines and starts mid-stream with Fleet Server runtime output, not the beginning of the E2E command. The visible tail shows TestStandAloneRunningSuite passing, followed by a package-level FAIL, which means the actual failing test output likely occurred earlier and was not included in the provided artifact.

The PR file list is limited to internal/pkg/bulk/* and a changelog fragment; it does not change .buildkite/scripts/e2e_test.sh, magefile.go, or testing/e2e/*, so there is no direct evidence in the available data that this E2E failure was introduced by the PR’s bulk changes.

Evidence

  • Build: https://buildkite.com/elastic/fleet-server/builds/15126
  • Job/step: E2E Test (.buildkite/scripts/e2e_test.sh)
  • E2E command path: .buildkite/scripts/e2e_test.sh:16 runs mage test:e2e test:junitReport; magefile.go:2162 runs go test -v -timeout 30m -tags=e2e,... -count=1 -race -p 1 ./... from testing/.
  • Key log excerpt from the available artifact:
--- PASS: TestStandAloneRunningSuite (549.96s)
    --- PASS: TestStandAloneRunningSuite/TestOpAMPWithEDOTCollector (123.16s)
    --- PASS: TestStandAloneRunningSuite/TestOpAMPWithUpstreamCollector (300.36s)
    --- PASS: TestStandAloneRunningSuite/TestWithElasticsearchConnectionFailures (41.94s)
    --- PASS: TestStandAloneRunningSuite/TestWithElasticsearchConnectionFlakyness (32.51s)
    --- PASS: TestStandAloneRunningSuite/TestWithSecretFiles (1.01s)
FAIL
FAIL github.com/elastic/fleet-server/testing/e2e 1309.997s
Error: exit status 1

Verification

  • Checked the provided failure summary and the only local Buildkite log artifact.
  • Searched the artifact for panic, Error Trace, DATA RACE, race detected, test timed out, signal: killed, and context deadline; none were present.
  • Checked existing flaky-test issues. There is an open OpAMP flaky issue (#6590), but the current tail shows both visible OpAMP subtests passed, so the available evidence does not support attributing this failure to that known flake.

Follow-up

If the full log shows a specific failing suite, rerun this detective on that output; the current tail is insufficient to identify a defensible code or test fix.


What is this? | From workflow: PR Buildkite Detective

Give us feedback! React with 🚀 if perfect, 👍 if helpful, 👎 if not.

@blakerouse blakerouse left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These are big improvements, good work on this.

@michel-laterman
michel-laterman merged commit d774f03 into elastic:main Jun 17, 2026
10 checks passed
@michel-laterman
michel-laterman deleted the worktree-bulk-benchmarks branch June 17, 2026 16:21
@michel-laterman michel-laterman linked an issue Jun 22, 2026 that may be closed by this pull request
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backport-skip Skip notification from the automated backport with mergify enhancement New feature or request Team:Elastic-Agent-Control-Plane Label for the Agent Control Plane team

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fleet-server bulker improvements

2 participants