Repository navigation
Conversation
The doc offsets are currently being delta encoded into encoded as grouped vints. For some queries reading grouped vints is a relative expensive operations. This PR changes grouped vint encoding with a more simplistic bit packing method. Because this PR makes a format a new new codec version is added and a bwc test is added to TsdbDocValueBwcTests (and version2 of doc values format was forked in test code).
|
Hi @martijnvg, I've created a changelog YAML for you. |
|
Buildkite benchmark this with clickbench-columnar-mode please |
…ew codec that uses BITPACKING codec. Before the grouped vint or bitpacking compression was picked based on codec version, but that doesn't work well in mixed clusters. The new codec is now used based on index version and that should work well in mixed clusters regardless of whether shard recovery is happening.
…offsets_using_bitpacking
64bc91e to
7631d66
Compare
|
Buildkite benchmark this with clickbench-columnar-mode please |
|
Buildkite benchmark this with clickbench-columnar-mode please |
|
Pinging @elastic/es-storage-engine (Team:StorageEngine) |
|
I didn't finish reviewing this PR. Looks great so far, but I want to understand more about the indexing versioning before approving. |
Something that I forgot to point out is that with the third commit, the CI ran all tests, but even with codec version check the serverless upgrade tests fail: https://buildkite.com/elastic/elasticsearch-serverless-es-pr-check/builds/64030 This made me realize the problem with codec version and why we can't rely on it for bwc checks. |
| if (indexCreatedVersion.onOrAfter(IndexVersions.TIME_SERIES_DOC_VALUES_FORMAT_VERSION_3)) { | ||
| return new ES819Version3TSDBDocValuesFormat(useLargeBlockSize); | ||
| } else { | ||
| return useLargeBlockSize ? ES819TSDBDocValuesFormat.getInstance(true) : ES819TSDBDocValuesFormat.getInstance(false); |
There was a problem hiding this comment.
nit: ES819TSDBDocValuesFormat.getInstance(useLargeBlockSize)
…offsets_using_bitpacking
496847d to
9ebbe3c
Compare
|
Buildkite benchmark this with clickbench-columnar-mode please |
💚 Build Succeeded
This build ran two clickbench-columnar-mode benchmarks to evaluate performance impact of this PR. History
|
…ic#142772) The doc offsets are currently being delta encoded into encoded as grouped vints. For some queries reading grouped vints is a relative expensive operations. This PR changes grouped vint encoding with a more simplistic bit packing method. Because this PR makes a format a new doc values format is added that extends from ES819TSDBDocValuesFormat. Only new indices will use this new doc value format. This is required because otherwise in mixed clusters bwc doc values format issues will still occur (even with a new codec version). This is because there is no mechanism that prevents shard recovery from a newer node with higher codec version to an older node with lower codec version. Tying the new doc values format to index version avoids this problem.
The doc offsets are currently being delta encoded into encoded as grouped vints. For some queries reading grouped vints is a relative expensive operations. This PR changes grouped vint encoding with a more simplistic bit packing method.
This change results in slightly lower disk usage and lower query latencies in cases binary doc values is used. In case of clickbench benchmark, a number of queries showed lower latencies: q28 from 16s to 13s, q27 from 715 ms to 551 ms, q39 from 1065ms to 802 ms.
Because this PR makes a format a new doc values format is added that extends from
ES819TSDBDocValuesFormat. Only new indices will use this new doc value format. This is required because otherwise in mixed clusters bwc doc values format issues will still occur (even with a new codec version). This is because there is no mechanism that prevents shard recovery from a newer node with higher codec version to an older node with lower codec version. Tying the new doc values format to index version avoids this problem.