Skip to content

[TSDB++] Facilitate backfilling in TSDB #149856

Description

@gmarouli

The problem

The way Time Series Data Streams (TSDS) work today is that after creating the initial index, users can not load historical data at any point in time, only data going forward from the start time that index supports.

The reason for that is that Time Series Data Streams (TSDS) use time bound backing indices that accept documents based on their timestamp values. A time bound index has a start and end time and accepts documents whose @timestamp belongs to that timeframe, anything else gets rejected. When you add a document to a TSDS, Elasticsearch adds the document to the appropriate backing index based on its @timestamp value.

As time passes, new time bound indexes are added to cover new time intervals. So, all time intervals are supported after the creation of a TSDS, but no data can be added before the initial start time as there is no way to add time bound indexes in the past.

This blocks non greenfield Elasticsearch users with existing time series data that want to migrate to using Elasticsearch TSDS. The most common use case is to start a new index with recent data, to either test how TSDS work or just adopt the feature, and start backfilling historical data at a later point or in the background.

Proposed solution

For green field use-cases

The proposed solution enables a TSDB data stream to lazily create backing indices in the past of predefined duration, potentially configurable via a cluster setting. The creation of the backing indices is triggered only when a document with a past @timestamp arrives and does not match any of the current backing indices while the @timestamp falls within the eligible write window.

The eligible write window is defined by the lifecycle feature managing a TSDB, specifically, either by the time condition of the first read-only action or the retention. For example, if TSDB should be downsampled after 7 days, this means that we will create new backing indices only for documents whose @timestamp falls within now till 7 days ago.

For historical data of longer periods

Considering the above limitation, we want to offer an alternative path for users who want to load historical data further than the write window. For them, we suggest the following steps:

  • Scale their target indexing tier to hold the whole amount of their historical data next to their current ingestion.
  • Create a new TSDS, let's call it historical-ds using a similar index template as the original one, potentially adjust shards if deemed necessary.
  • Start indexing to that historical-ds only the historical data and keep indexing current data to the original one.
  • When finished, enable DLM if it's not enabled and add any downsampling, frozen or retention configurations.
  • Remember to delete the historical-ds when retention the data "expires".

Caveats

This scenario has a few caveats:

  • Queries need to be adjusted to include both original & historical data streams, or the user needs to define a data stream alias.
  • When DLM is enabled on a data stream with a lot of data it can trigger a large amount of heavy operations that can overwhelm the master. This is already an issue, but with this proposal we make it more prominent. We would like to add a throttling mechanism to "pace" at downsampling and searchable snapshot operation per data stream.
  • This scenario is supported by DLM only. We cannot support such throttling in ILM. The user can manually add policies per index if they wish to, but we will not provide any native ILM throttling mechanism.

Planned work

Activity

  1. added theissue type on May 26, 2026
  2. added and removed
    needs:triageRequires assignment of a team area label
    on May 26, 2026
  3. elasticsearchmachine commented on May 26, 2026

    @elasticsearchmachine
    Collaborator

    Pinging @elastic/es-storage-engine (Team:StorageEngine)

  4. gustavo89587 commented on Jun 15, 2026

    @gustavo89587

    Have you considered introducing weighted fairness instead of per-stream throttling only?
    A single historical backfill operation on a large TSDS could monopolize DLM execution slots and delay smaller streams indefinitely. Even simple aging or priority escalation mechanisms might help prevent starvation while preserving cluster stability.
    Curious if this trade-off has already been evaluated.

  5. jbaiera commented on Jun 15, 2026

    @jbaiera
    Contributor

    could monopolize DLM execution slots and delay smaller streams indefinitely

    @gustavo89587 I'm not sure what you mean by DLM execution slots - DLM executes for all data streams with attached lifecycles in a cluster per polling event. The limitation is per individual data stream. A large data stream with historical indices being throttled in downsampling will have no effect on any smaller streams.

  6. gustavo89587 commented on Jun 15, 2026

    @gustavo89587

    @jbaiera The fallacy of this data stream segregation ignores the low-level topology. The problem is not the logical limit of the stream, but the saturation of the transport thread pools and the contention on the node heap during massive I/O operations. While the ILM processes the backfilling, the node's queue depth locks the scheduler. If multiple backfills run, cache thrashing and IOPS competition destroy the performance of any priority stream. The proposed design forgets that hardware does not compartmentalize I/O in the same way that software segments streams. Without a control that acts on the node's real backpressure, considering the state of the thread pools, this "limitation" is merely a palliative that fails under real load.

  7. gustavo89587 commented on Jun 15, 2026

    @gustavo89587

    I understand the focus on logical segregation by data stream, but I see a critical blind spot here: resource contention at the node level. In Elasticsearch, backfilling management doesn't happen in a vacuum; ILM, downsampling operations, and new data indexing share the same transport threads and the indexing thread pool.

    When backfilling triggers without a backpressure mechanism monitoring the overall health of the queue depth and heap consumption, we risk massive history operations saturating I/O and causing page cache thrashing, degrading the performance of priority streams. Stream throttling is just access control, not a QoS policy.

    To make this safe in production environments, I suggest we look at:

    Implementing token-based weighted fairness for DLM tasks, ensuring that smaller streams are not choked by the backfilling latency of a massive stream.

    Active monitoring of the transport thread pool queue depth as a dynamic throttle trigger, ensuring that the system slows down the backfill rate before the node starts rejecting legitimate requests due to thread pool exhaustion.

    If the idea is to maintain the current implementation, what is the planned mechanism to prevent thread pool collapse during simultaneous backfill spikes?

  8. jbaiera commented on Jun 16, 2026

    @jbaiera
    Contributor

    @gustavo89587 I'm sorry, but these comments are nonsense. Given your other contributions to the project so far (and to other projects on Github for that matter) I have to assume these posts are generated. If you have primary concerns about data management and indexing resource management in Elasticsearch, then I suggest you raise them in a different setting than this issue as they are clearly beyond the scope of the work being discussed here.

  9. gustavo89587 commented on Jun 16, 2026

    @gustavo89587

    @jbaiera I use AI as a tool to structure my reasoning, like any engineer does, and I guarantee you do too, even if you don't say so and criticize me... the ideas are mine, AI doesn't generate what's in my head, LOL. I may be wrong on technical aspects, and tell me... who doesn't make mistakes?? But I'm here to learn and genuinely contribute....

  10. jbaiera commented on Jun 16, 2026

    @jbaiera
    Contributor

    @gustavo89587 Again, if you have concerns about resource management in Elasticsearch, then I suggest you raise them in a different setting than this issue as they are beyond the scope of the work being discussed here.

  11. gmarouli commented on Jun 16, 2026

    @gmarouli
    ContributorAuthor

    @gustavo89587 Again, if you have concerns about resource management in Elasticsearch, then I suggest you raise them in a different setting than this issue as they are beyond the scope of the work being discussed here.

    I second this, currently, elasticsearch treats all data the same, different data streams and indices are indexed and managed as soon as possible. The same applies here. This projects enables users to index data beyond the start time of a TSDB data stream making no assumptions on how important the data is. The user is responsible to scale up or manage the rate of ingesting data (historical or current) based on their priorities.

    If there is a ever a different solution, we can always evaluate how it could work along this feature.

  12. sosmani15 commented on Oct 5, 2026

    @sosmani15
    Contributor

    @gmarouli can you please update the status of this ticket and close it if there's no more work pending?

  13. gmarouli commented on Oct 6, 2026

    @gmarouli
    ContributorAuthor

    @sosmani15 the ticket is up to date, the docs are not merged yet: elastic/docs-content#7252

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions