There have been cases reported of relatively small Logstash pipelines taking a long time to load, or simply not loading up at all, in the Pipeline Viewer Monitoring UI.
An easy way to test this would be to create a test Logstash pipeline containing ~30 vertices, including a few if vertices (so we have > ~30 edges). Configure Logstash Monitoring and let the pipeline run for several hours (I'll update the issue with more concrete numbers based on my testing). Then open the Pipeline Viewer Monitoring UI for the pipeline.
The Kibana API that the Pipeline Viewer calls for its data returns the entire structure of the pipeline (vertices + edges, see the pipeline.representation.graph property in the API response). Each vertex object has a stats property which is an object representing stats about that vertex. There can be up to 7 stats for each vertex. Each stat object has a data property which contains timeseries data for that stat (an array of timestamp+metric pairs).
All in all, this generates a LOT of data points. A pipeline with 30 vertices would generate 30 x 7 x (# of timeseries buckets) data points. The # of timeseries buckets is determined by how long the pipeline version being viewed has been running. For instance, if it has been running for 24 hours, that would set the date_histogram bucket size to 10 minutes, resulting in 144 timeseries buckets (24 * 60 / 10 = 144). So in the 30-vertex example, that would result in a total of 30 x 7 x 144 = 30,240 data points!
However, if you look at the pipeline's rendering in the Pipeline Viewer, you'll only see 3 stats are shown per vertex. The remaining stats aren't visible but they are used for calculations that could affect the design of the vertex (e.g. highlighting it if the plugin represented by the vertex is consuming an unusually high amount of CPU).
Also, only the latest data point from the timeseries data for each of the 7 stats is considered for calculations or display. So why do we even get a timeseries from ES? Why not just get the latest data point? That's because if you click on a vertex, a vertex detail flyout opens up to the right. In this flyout, the timeseries for some of the stats is rendered as sparklines.
I suspect the performance issues caused by this page are related to large number of data points that the ES aggregation query has to generate (but we should confirm this with testing). If this is indeed the case, here might be one way to mitigate the issue:
For the initial rendering of the pipeline (i.e. with the detail flyout closed), only request the latest data points for each stat instead of a timeseries for each stat from ES (as opposed to today, where we request the timeseries for each stat, then pull out the latest data point from the timneseries in browser-side code). Then, when the user clicks on a particular vertex, request the timeseries for each stats for just that vertex, in order to render the sparklines in the detail flyout.
There have been cases reported of relatively small Logstash pipelines taking a long time to load, or simply not loading up at all, in the Pipeline Viewer Monitoring UI.
An easy way to test this would be to create a test Logstash pipeline containing ~30 vertices, including a few
ifvertices (so we have > ~30 edges). Configure Logstash Monitoring and let the pipeline run for several hours (I'll update the issue with more concrete numbers based on my testing). Then open the Pipeline Viewer Monitoring UI for the pipeline.The Kibana API that the Pipeline Viewer calls for its data returns the entire structure of the pipeline (vertices + edges, see the
pipeline.representation.graphproperty in the API response). Each vertex object has astatsproperty which is an object representing stats about that vertex. There can be up to 7 stats for each vertex. Each stat object has adataproperty which contains timeseries data for that stat (an array of timestamp+metric pairs).All in all, this generates a LOT of data points. A pipeline with 30 vertices would generate 30 x 7 x (# of timeseries buckets) data points. The # of timeseries buckets is determined by how long the pipeline version being viewed has been running. For instance, if it has been running for 24 hours, that would set the
date_histogrambucket size to 10 minutes, resulting in 144 timeseries buckets (24 * 60 / 10 = 144). So in the 30-vertex example, that would result in a total of 30 x 7 x 144 = 30,240 data points!However, if you look at the pipeline's rendering in the Pipeline Viewer, you'll only see 3 stats are shown per vertex. The remaining stats aren't visible but they are used for calculations that could affect the design of the vertex (e.g. highlighting it if the plugin represented by the vertex is consuming an unusually high amount of CPU).
Also, only the latest data point from the timeseries data for each of the 7 stats is considered for calculations or display. So why do we even get a timeseries from ES? Why not just get the latest data point? That's because if you click on a vertex, a vertex detail flyout opens up to the right. In this flyout, the timeseries for some of the stats is rendered as sparklines.
I suspect the performance issues caused by this page are related to large number of data points that the ES aggregation query has to generate (but we should confirm this with testing). If this is indeed the case, here might be one way to mitigate the issue:
For the initial rendering of the pipeline (i.e. with the detail flyout closed), only request the latest data points for each stat instead of a timeseries for each stat from ES (as opposed to today, where we request the timeseries for each stat, then pull out the latest data point from the timneseries in browser-side code). Then, when the user clicks on a particular vertex, request the timeseries for each stats for just that vertex, in order to render the sparklines in the detail flyout.