Repository navigation
Improve loading time of Kubernetes package Dashboards #31021
Description
Activity
- addedTeam:Cloudnative-MonitoringLabel for the Cloud Native Monitoring teamLabel for the Cloud Native Monitoring team
on Mar 28, 2022 - changed the title
[-]Optimise Kubernetes Dashboards[/-][+]Optimize Kubernetes Dashboards[/+]on Mar 28, 2022 MichaelKatsoulis commented
on Mar 30, 2022 ContributorAuthorMore actionsMy suggestion regarding the default kubernetes dashboard optimization is to split it into 2 different dashboards.
The split can be based on the concept of each visualisation's data.
Meaning that some of them make sense to be displayed per time, while others make more sense to display the current value as a number.For example for the number of available/desired/unavailable pods or number of nodes it is most important is to display the current situation in the cluster.
While for other visualisations like
Top CPU intensive podsorCPU utilization per nodeit would be insightful to display the evolution of the value per time. A user would like to see how the cpu utilisations of a specific pod or node has changed over the past week.The two dashboards could like this:
Time series dashboard:
This grouping can make the dashboards more performant as less queries will be performed simultaneously.
Also costly queries with aggregations over big time range will only be performed for the vis that make sense.Reacted by Giuseppe SantoroMichaelKatsoulis commented
on Mar 30, 2022 ContributorAuthorMore actionscc @ChrsMark , @tetianakravchenko , @ruflin, @mlunadia, @gsantoro
++ on splitting up the dashboards. Will the dashboards link to each other?
That would be great. We do the same for Istio module to split the control plane from data plane views:
https://www.elastic.co/guide/en/beats/metricbeat/current/metricbeat-module-istio.html#_dashboard_30MichaelKatsoulis commented
on Mar 31, 2022 ContributorAuthorMore actionsYes I was thinking something like Istio tab view! That would be great! Does this allow to set different time ranges to each one?
I thought by now there is a new / better way on how to link dashboards together. @alexfrancoeur You might be able to point us to the right direction here?
Thanks for the ping @ruflin. I think there are a number of best practices these integration dashboards can start to leverage. I've listed a bunch here in the past.
Kibana has drilldown capabilities in a dashboard (https://www.elastic.co/guide/en/kibana/current/drilldowns.html#create-drilldowns). This is great for creating workflows from dashboard to dashboard. An overview dashboard to a details dashboard for example. We support dashboard to dashboard and dashboard to external URLs (paid feature). For the integrations dashboard, a combination of markdown for general navigation and drilldowns for workflows is probably the best option.
If you'd like to sit down with some kibana folks and discuss best practices, we're happy to engage. For example we no longer need to ship 100's of visualizations referenced by a dashboard, we can simplify to package all in a single dashboard JSON now.
This is off topic, but while we're on the topic of dashboard linking, I think it's worth raising if we should be linking to solutions as well. There is probably some low hanging fruit here. Rather than taking that context and navigating to a dashboard, we could apply it to a solution view to create solution drilldowns. I hack together this all the time for demos using the URL drilldowns. Meaning if there's a host IP in a dashboard, let's click into that to navigate to a filtered view in the metrics app. Building these experiences as part of our integrations add for a much more integrated experience when onboarding a new data source. If there's interest in collaborating on something like this, let's have a quick chat with myself and @sixstringcode
Thanks for the list @alexfrancoeur . The drilldown one is the one I was looking for. There lots of other great hints in the issue you linked.
About "as value" I just had a conversation with your team and we should figure out ways how to automatically convert it. @ChrsMark @MichaelKatsoulis If we redo the k8s dashboards, lets use these best practices directly as an example and also switch to "value".
On the linking to solutions, ++. It is a topic we should also involve @jasonrhodes from unified observability.
Reacted by Christos Markou, Michalis Katsoulis and AlexFI've been attending the TSDB meeting yesterday and @Mpdreamz showed off a demo for using TSDB in APM. A first version of the metrics parts are merged into main in Elasticsearch and available in the snapshot builds. I think it is also worth trying this out for the k8s data to see what impact it has.
My understanding is that currently the storage and query part are available but we can't make use of it yet in Kibana. Also in the package-spec we don't support the time series fields yet. What it means is that we have to adjust the templates manually and see what affect it has.
@imotov Tried to find some public docs I can point the team to around mappings and TSDI but was not successful. Is this already available?
Reacted by Christos MarkouReacted by Christos MarkouI filed elastic/package-spec#311 to get support for TSDB in the package-spec.
MichaelKatsoulis commented
on Apr 1, 2022 ContributorAuthorMore actionsI read in the best practises and by @ruflin suggestions that moving to lens is the way forward. I don't see or maybe I don't know how some of our tsvb visualisations can be moved to lens.
I will give an example regarding
desired pods. The fieldkubernetes.deployment.replicas.desiredhas a value per deployment likekube_deployment_spec_replicas{namespace="default",deployment="hello-python"} 1 kube_deployment_spec_replicas{namespace="kube-system",deployment="coredns"} 2 kube_deployment_spec_replicas{namespace="kube-system",deployment="kube-state-metrics"} 1 kube_deployment_spec_replicas{namespace="local-path-storage",deployment="local-path-provisioner"} 1and we want to sum up all the last values of this fields for all deployments.
If we compare seeming the same dashboards with same query in tsvb and lens we can spot huge differences.None of the results is the correct one. But lens one extreme!
Tsvb result is actually affected by the interval.There where discussions about this in elastic/integrations#2159 (comment) and @ChrsMark updated the tsvb query by using series aggregation and grouping by deployment name.
But as long as this is not supported in Lens, I don't see how we can use it for such cases.
5 remaining items
Completely agree, that’s what we will work on for Lens in 8.3
Reacted by Michalis KatsoulisMichaelKatsoulis commented
on Apr 7, 2022 ContributorAuthorMore actionsAn extra thing that could be discussed regarding drilldown is the user experience. Instead of the user having to press the options button and then select the drill down name like:
There should be an easier and more clear way.
If I were the user I would not understand that this red1on the vis means that there is a drilldown, and in order to see it I need two more steps.
Probably pressing on the1(or whatever that makes more sense) should navigate them to the dash.- The
1is only visible in edit mode, it's not shown in view mode (which should be the common case for users) - The easiest integration for [Lens] Allow metric visualization to drill down kibana#122879 is to allow the user to click into the visualization (e.g. on the "pods" text), then getting a context menu which allows them to navigate. We can think about how to provide an affordance during implementation
- The
MichaelKatsoulis commented
on Apr 7, 2022 ContributorAuthorMore actions@flash1293 I agree with your second bullet. That would be a good way. As it is now, there is no way a user can understand there is something more. The word
drilldownalso does not mean anything to someone that doesn't know what it is.MichaelKatsoulis commented
on Apr 7, 2022 ContributorAuthorMore actionsMichaelKatsoulis commented
on Apr 18, 2022 ContributorAuthorMore actions@flash1293 Could we arrange a zoom call whenever possible to ask you about best ways to show some metrics in Kibana?
I want to create some nice gauges but to get those numbers, series aggregations are needed and then mathematical formulas like division.
I can do things like this. But cannot get to use those two number for a division to get the percentage.
@katefarrar @mlunadia it would be great for us to try to understand what it is about these dashboards that people want/need/use as we try to think through the infrastructure UI.
Reacted by Kate FarrarMichaelKatsoulis commented
on Apr 21, 2022 ContributorAuthorMore actionsWe had a nice discussion with @flash1293 about ways to create some visualizations and we concluded that some things are not possible yet. But they can be in the near future.
Until then we can use some workarounds when showing informations likememory reserved,memory used,cores reserved,cores used,pods reservedusing the mark-down option.
Ideally we would like to be able to math calculation with the numbers in each vis to get the percentage.
Also we are waiting for the drill down option to be available in Lens in 8.3 or 8.4 release to better connect the dashboards between each other.@jasonrhodes 100% we have plans to tackle this holistically and will for now address any low hanging fruit. We have already started working on establishing a baseline with different discovery activities one of them will be bringing your input in.
Reacted by Jason Rhodes@elastic/infra-monitoring-ui I wonder if there are things we can learn from the optimizations in this issue that could be applied to any other querying we are doing for infra UI.
I guess there are two things we can do:
-
From a joint product perspective look at which visualizations we have in the UI today that could be changed to a gauge or single value instead of a trend line. I think today we almost only use trend lines? Changing that could allow us to change to a more performant query while also giving better feedback about the data to the user. Having multiple types of visualizations feels natural.
-
Optimize the queries themselves. I'm a bit hesitant about this since it also requires the in-depth domain knowledge about which field means what and which aggregations causes that field to mean something else. What we could do however is try to take stock of the queries we do and how they perform (similar to the SM work we're doing) and then pick the top X and see if we can optimize them or feed them into point 1.
-
@miltonhultgren thanks! These sound like good ideas to me. I'm wondering if there are specific optimizations made in the work related to this issue (from @MichaelKatsoulis and others) that we could use to inform how we might optimize our own queries, but you're right that there is likely some work we'll need to do to understand whether that kind of overlap exists.
- changed the title
[-]Optimize Kubernetes Dashboards[/-][+]Improve loading time of Kubernetes package Dashboards [/+]on Apr 28, 2022 Further improvements will take place with the usage of TSDB features. Investigations will take place along with Rally framework.
We have found some opportunities for optimization in the dashboards::
Top CPU intensive pods gets the max of
kubernetes.container.cpu.usage.core.ns, then uses derivative aggregation over it and keeps the positive values. We can simply usekubernetes.container.cpu.usage.node.pctinstead and group by the pod name.Same for Top Memory intensive pods
CPU Usage by node sums all cpu usage nanocores per container, then uses a painless script to normalise it to the metricset period and groups by the node name. Instead we can use the node metric
Kubernetes.node.cpu.usage.nanocoresand divide it withkubernetes.node.cpu.allocatable.cores. Same approach is used in metrics UISame for Memory Usage by node. We can divide
kubernetes.node.memory.usage.bytestokubernetes.node.memory.allocatable.bytesSame approach for network in and out bytes
I tested that by creating a separate dashboard with all those visualisations optimised and the loading time for 24h range decreased from 1m and 10 seconds down to 30 seconds.