Skip to content

[meta, 8.5, 8.6, 8.7] Logstash Pipeline Health API - Pipeline Flow Metrics #14463

Description

@yaauie

Problem Statement

It is often difficult to understand the health of a pipeline, including whether it is exerting or propagating back-pressure or otherwise staying reasonably “caught up” with its inputs. The cumulative-value metrics exposed by the API are not typically useful on their own, and when we are debugging a pipeline we often require a pair of captures and manual math in order to see the flow of events through the pipeline.

We want to provide better visibility into the flow of events through a Logstash pipeline in order to easier pinpoint the source of degraded status.

Scope and Goals

We will provide visibility into the current, recent, and lifetime flow rates of pipeline-level metrics in the /_node/stats API, so that our users can have more contextual information when analyzing the throughput of a pipeline.

We will do so in a way that:

  • (a) provides a prototype for future efforts to also supply plugin-level metrics, and
  • (b) allows us to use these newly-available flow metrics in a future effort to establish opinionated and configurable heuristics for upward-propagating plugin, pipeline, and node health.

Design

We implement a phased approach, with each phase designed to produce some quantity of value while unlocking the development of one or more future phases. Because each phase is designed to be forward-compatible with the phases that follow it, we are not bound to ship this entire effort in a single Logstash release.

Phase 0: Collecting Flow Metrics

Status: included in 8.5.0

We build a metric collector that periodically queries a subset of metrics we care about for each pipeline, capturing the current value into an internal data structure.

The query cadence must be regular and configurable, with an initial target of every 1s, so that we can tune the cadence using user feedback during this feature's Technology Preview phase. Our code will log a warning when it fails to run as frequently as configured.

For each pipeline, we should capture:

  • input_throughput: pipelines.*.events.in / second of pipeline uptime
  • filter_throughput: pipelines.*.events.filtered / second of pipeline uptime
  • output_throughput: pipelines.*.events.out / second of pipeline uptime
  • queue_backpressure: pipelines.*.queue_push_duration_in_millis / millisecond of pipeline uptime

    NOTE: positive numbers indicate the presence of back-pressure, but the scale is determined by the quantity and type of inputs and is not comparable from one pipeline to another.

  • worker_concurrency: duration (milliseconds spent working) / millisecond of pipeline uptime

    NOTE: both numerator and denominator are milliseconds, so this number is unitless.

At a minimum, the data-structure we capture into MUST:

  • be capable of prducing "current" rate using the most recent two captures,
  • be capable of producing a "lifetime" rate using the oldest and most-recent captures,
  • be concurrency-safe, and
  • adhere the following pseudo-interface:
    package org.logstash.instrument.metrics;
    
    interface FlowMetric {
        void capture()
        String getName()
        Map<String,Number> getValue()
    }
    
    interface FlowMetricFactory {
        FlowMetric create(String name, Metric numerator, Metric Denominator)
    }

The data-structure also SHOULD:

  • implement its concurrency-safety without the use of locks or blocking mechanisms.

In order to support the common denominator second of pipeline uptime, we will need to also introduce a new system-clock-backed metric whose persistence across pipeline reloads matches the persistence of the other metrics (that is: since our numerator metrics are cleared when a pipeline is reloaded, then the denominator should also be cleared).

Phase 1.A: Presenting Flow Metrics in Node Stats API

Status: included in 8.5.0

The Node Stats API is extended to produce rates in addition to the existing point-in-time cumulative value of the metrics, allowing a consumer to get an instantaneous view of the current flow-rates of the pipeline in a single API request. As the metrics identified in Phase 0 are all inside the pipeline.*.events namespace, these new metrics will be in a sibling flow namespace, and will each contain a key/value map containing all available flow rates.

For example,

  "pipelines" : {
    "my_pipeline" : {
      "events" : {
        "duration_in_millis" : 56493,
        "out" : 33745,
        "in" : 33745,
        "queue_push_duration_in_millis" : 188,
        "filtered" : 33745
      },
      "flow": {
        "worker_concurrency": { "current": 3.87, "lifetime": 2.12 },
        "input_throughput": { "current": 2872.1, "lifetime": 1889.7 },
        "filter_throughput": { "current": 2792.4, "lifetime": 1701.4 },
        "output_throughput": { "current": 2801.7, "lifetime": 1701.3 },
        "queue_backpressure": { "current": 0, "lifetime": 0.1 },
      }

Flow rate numerators and denominators that are time-based will be normalized to "seconds".

Phase 2.A: Extended Flow Metrics

Status: included in 8.6.0

The current near-instantaneous rate of a metric is not always useful on its own, and a lifetime rate is often not granular enough for long-running pipelines. These rates becomes more valuable when they can be compared against rolling average rates for fixed time periods, such as "the last minute", or "the last day", which allows us to see that a degraded status is either escalating or resolving.

We extend our historic data structure to retain sufficient data-points to produce several additional flow rates, limiting the retention granularity of each of these rates at capture time to prevent unnecessary memory usage. This enables us to capture frequently enough to have near-instantaneous "current" rates, and to produce accurate rates for larger reasonably-accurate windows.

EXAMPLE: A last_15_minutes rate is calculated using the most-recent capture, and the youngest capture that is older than roughly 15 minutes, by dividing the delta between their numerators by the delta between their numerators. If we are capturing every 1s and have a window tolerance of 30s, the data-structure compacts out entries at insertion time to ensure it is retaining only as many data-points as necessary to ensure at least one entry per 30 second window, allowing a 15-minute-and-30-second window to reasonably represent the 15-minute window that we are tracking (retain 90 of 900 captures). Similarly a 24h rate with 15m granularity would retain only enough data-points to ensure at least one entry per 15 minutes (retain ~96 of ~86,400 captures).

In Technology Preview, we will with the following rates without exposing configuration to the user:

  • last_1_minute: 1m@3s
  • last_5_minutes: 5m@15s
  • last_15_minutes: 15m@30s
  • last_1_hour: 1h @ 60s
  • last_24_hours: 24h@15m

These rates will be made available IFF their period has been satisfied (that is: we will not include a rate for last_1_hour until we have 1 hour's worth of captures for the relevant flow metric).
Additionally with the above framework in place, our existing current metric will become as an alias for 10s@1s, ensuring a 10-second window to reduce jitter.

Phase 2.B: Persistent Queue Metrics

Status: included in 8.6.0

We know that A Persistent Queue that is consistently growing is one that will eventually run out of capacity and "overflow" into back-pressure to the pipeline's inputs. By adding flow metrics about the queue's unread events and usage on disk, we can help our users discover issues before they become a problem.

PQ metrics metrics will be supplied inside the pipeline.*.flow node along-side the other pipeline-level flow metrics, prefixed with queue_persisted_* so as to provide leading context scoping them to a pipeline's queue that is of a persisted type.

  • queue_persisted_growth_events: pipeline.*.queue.events_count / second of pipeline uptime : the rate of change of the PQ's unacked events should trend toward zero (or negative), indicating that we are processing events as quickly as they are being put into the queue.
  • queue_persisted_growth_bytes: pipeline.*.queue.queue_size_in_bytes / second of pipeline uptime : this one is liable to be jumpy, as the numerator includes events that have been acked on pages that haven't been fully-acked and expunged. Long-range trend toward zero or negative is ideal.

Future Development (Out of Scope)

Plugin-level Metrics

Status: included in 8.7.0

We will use the pipeline-level metrics as a prototype for plugin-level metrics, so that users can dig into the cause of degraded status, whether it is an output that is experiencing back-pressure resulting in a throughput drop, a filter that is taking more CPU-time per event than usual, an input that is not receiving events from upstream, etc.

Because the filter-and-output phase of a pipeline is typically run in parallel, filters and outputs will likely benefit from having their metric rates against both the wall-clock and the CPU clock.

The volume of per-plugin captures may force us to tune the collector developed in Phase 0 to run concurrently or attempt a less-frequent cadence.

Traffic-light Status Heuristics

We would like to get to a place where we can opinionatedly tell a user that either they don't have to worry about a pipeline (green), that they may need to dig in deeper soon (yellow), or that urgent intervention is required (red). Having historic metric flow-rates is going to be a big part of empowering opinionated heuristics. This effort is meant as a prerequisite for this future effort, and to provide the data with which to make such heuristics meaningful.

Implementation Plan

Activity

  1. changed the title [-]META: Pipeline Flow Metrics[/-] [+][meta] Pipeline Flow Metrics[/+] on Aug 24, 2022
  2. changed the title [-][meta] Pipeline Flow Metrics[/-] [+][meta] Pipeline Health API - Pipeline Flow Metrics[/+] on Aug 24, 2022
  3. changed the title [-][meta] Pipeline Health API - Pipeline Flow Metrics[/-] [+][meta, 8.5, 8.6] Pipeline Health API - Pipeline Flow Metrics[/+] on Sep 19, 2022
  4. changed the title [-][meta, 8.5, 8.6] Pipeline Health API - Pipeline Flow Metrics[/-] [+][meta, 8.5, 8.6] Logstash Pipeline Health API - Pipeline Flow Metrics[/+] on Sep 19, 2022
  5. mashhurs commented on Sep 19, 2022

    @mashhurs
    Contributor

    This PR covers Phase 0 and Phase 1.A.

  6. deleted a comment from mashhurs on Oct 19, 2022
  7. roaksoax commented on Jan 27, 2023

    @roaksoax
    Contributor

    Provided the the bulk of the work is now complete, I'm going to go ahead an close this issue. The Traffic-Ligh status heuristics seats outside of Pipeline Flow Metrics.

  8. changed the title [-][meta, 8.5, 8.6] Logstash Pipeline Health API - Pipeline Flow Metrics[/-] [+][meta, 8.5, 8.6, 8.7] Logstash Pipeline Health API - Pipeline Flow Metrics[/+] on Jun 8, 2023
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions