Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to AIOps Using Open Source Solutions
- Core AIOps concepts and their advantages
- The role of Prometheus and Grafana in the observability stack
- The place of ML in AIOps: predictive versus reactive analysis
Configuring Prometheus and Grafana
- Installing and setting up Prometheus for time series data collection
- Building Grafana dashboards using live metrics
- Managing exporters, relabeling, and service discovery
Preparing Data for Machine Learning
- Extracting and processing Prometheus metrics
- Curating datasets suitable for anomaly detection and forecasting
- Applying Grafana transformations or Python-based data pipelines
Utilizing Machine Learning for Anomaly Detection
- Foundational ML models for outlier identification (e.g., Isolation Forest, One-Class SVM)
- Training and assessing models against time series data
- Displaying detected anomalies within Grafana dashboards
Forecasting Metrics via Machine Learning
- Developing forecasting models (ARIMA, Prophet, and an introduction to LSTM)
- Anticipating system load and resource consumption
- Leveraging predictions for proactive alerting and scaling strategies
Combining ML with Alerting and Automation
- Establishing alert rules driven by ML insights or static thresholds
- Implementing Alertmanager and notification routing
- Initiating scripts or automation workflows upon anomaly detection
Scaling and Implementing AIOps in Production
- Connecting external observability tools (e.g., ELK stack, Moogsoft, Dynatrace)
- Operationalizing ML models within observability workflows
- Best practices for scaling AIOps implementations
Recap and Future Pathways
Requirements
- A solid grasp of system monitoring and observability fundamentals.
- Practical experience with Grafana or Prometheus.
- Proficiency in Python and an understanding of basic machine learning concepts.
Target Audience
- Observability engineers.
- Infrastructure and DevOps teams.
- Monitoring platform architects and Site Reliability Engineers (SREs).
14 Hours