Repository navigation
[TSDB++] Facilitate backfilling in TSDB #149856
Description
Activity
- addedneeds:triageRequires assignment of a team area labelRequires assignment of a team area label
on May 26, 2026 - added:StorageEngine/TSDBYou know, for MetricsYou know, for Metricsand removedneeds:triageRequires assignment of a team area labelRequires assignment of a team area label
on May 26, 2026 elasticsearchmachine commented
on May 26, 2026 CollaboratorMore actionsPinging @elastic/es-storage-engine (Team:StorageEngine)
Have you considered introducing weighted fairness instead of per-stream throttling only?
A single historical backfill operation on a large TSDS could monopolize DLM execution slots and delay smaller streams indefinitely. Even simple aging or priority escalation mechanisms might help prevent starvation while preserving cluster stability.
Curious if this trade-off has already been evaluated.could monopolize DLM execution slots and delay smaller streams indefinitely
@gustavo89587 I'm not sure what you mean by DLM execution slots - DLM executes for all data streams with attached lifecycles in a cluster per polling event. The limitation is per individual data stream. A large data stream with historical indices being throttled in downsampling will have no effect on any smaller streams.
@jbaiera The fallacy of this data stream segregation ignores the low-level topology. The problem is not the logical limit of the stream, but the saturation of the transport thread pools and the contention on the node heap during massive I/O operations. While the ILM processes the backfilling, the node's queue depth locks the scheduler. If multiple backfills run, cache thrashing and IOPS competition destroy the performance of any priority stream. The proposed design forgets that hardware does not compartmentalize I/O in the same way that software segments streams. Without a control that acts on the node's real backpressure, considering the state of the thread pools, this "limitation" is merely a palliative that fails under real load.
I understand the focus on logical segregation by data stream, but I see a critical blind spot here: resource contention at the node level. In Elasticsearch, backfilling management doesn't happen in a vacuum; ILM, downsampling operations, and new data indexing share the same transport threads and the indexing thread pool.
When backfilling triggers without a backpressure mechanism monitoring the overall health of the queue depth and heap consumption, we risk massive history operations saturating I/O and causing page cache thrashing, degrading the performance of priority streams. Stream throttling is just access control, not a QoS policy.
To make this safe in production environments, I suggest we look at:
Implementing token-based weighted fairness for DLM tasks, ensuring that smaller streams are not choked by the backfilling latency of a massive stream.
Active monitoring of the transport thread pool queue depth as a dynamic throttle trigger, ensuring that the system slows down the backfill rate before the node starts rejecting legitimate requests due to thread pool exhaustion.
If the idea is to maintain the current implementation, what is the planned mechanism to prevent thread pool collapse during simultaneous backfill spikes?
@gustavo89587 I'm sorry, but these comments are nonsense. Given your other contributions to the project so far (and to other projects on Github for that matter) I have to assume these posts are generated. If you have primary concerns about data management and indexing resource management in Elasticsearch, then I suggest you raise them in a different setting than this issue as they are clearly beyond the scope of the work being discussed here.
@jbaiera I use AI as a tool to structure my reasoning, like any engineer does, and I guarantee you do too, even if you don't say so and criticize me... the ideas are mine, AI doesn't generate what's in my head, LOL. I may be wrong on technical aspects, and tell me... who doesn't make mistakes?? But I'm here to learn and genuinely contribute....
@gustavo89587 Again, if you have concerns about resource management in Elasticsearch, then I suggest you raise them in a different setting than this issue as they are beyond the scope of the work being discussed here.
@gustavo89587 Again, if you have concerns about resource management in Elasticsearch, then I suggest you raise them in a different setting than this issue as they are beyond the scope of the work being discussed here.
I second this, currently, elasticsearch treats all data the same, different data streams and indices are indexed and managed as soon as possible. The same applies here. This projects enables users to index data beyond the start time of a TSDB data stream making no assumptions on how important the data is. The user is responsible to scale up or manage the rate of ingesting data (historical or current) based on their priorities.
If there is a ever a different solution, we can always evaluate how it could work along this feature.
@gmarouli can you please update the status of this ticket and close it if there's no more work pending?
@sosmani15 the ticket is up to date, the docs are not merged yet: elastic/docs-content#7252
The problem
The way Time Series Data Streams (TSDS) work today is that after creating the initial index, users can not load historical data at any point in time, only data going forward from the start time that index supports.
The reason for that is that Time Series Data Streams (TSDS) use time bound backing indices that accept documents based on their timestamp values. A time bound index has a start and end time and accepts documents whose @timestamp belongs to that timeframe, anything else gets rejected. When you add a document to a TSDS, Elasticsearch adds the document to the appropriate backing index based on its @timestamp value.
As time passes, new time bound indexes are added to cover new time intervals. So, all time intervals are supported after the creation of a TSDS, but no data can be added before the initial start time as there is no way to add time bound indexes in the past.
This blocks non greenfield Elasticsearch users with existing time series data that want to migrate to using Elasticsearch TSDS. The most common use case is to start a new index with recent data, to either test how TSDS work or just adopt the feature, and start backfilling historical data at a later point or in the background.
Proposed solution
For green field use-cases
The proposed solution enables a TSDB data stream to lazily create backing indices in the past of predefined duration, potentially configurable via a cluster setting. The creation of the backing indices is triggered only when a document with a past
@timestamparrives and does not match any of the current backing indices while the@timestampfalls within the eligible write window.The eligible write window is defined by the lifecycle feature managing a TSDB, specifically, either by the time condition of the first read-only action or the retention. For example, if TSDB should be downsampled after 7 days, this means that we will create new backing indices only for documents whose
@timestampfalls within now till 7 days ago.For historical data of longer periods
Considering the above limitation, we want to offer an alternative path for users who want to load historical data further than the write window. For them, we suggest the following steps:
historical-dsusing a similar index template as the original one, potentially adjust shards if deemed necessary.historical-dsonly the historical data and keep indexing current data to the original one.historical-dswhen retention the data "expires".Caveats
This scenario has a few caveats:
Planned work
- Detection in the transport bulk actions appears to cause no performance regression, so we are moving forward.
Add telemetry