Get in Touch

Course Outline

Introduction to AIOps Using Open Source Solutions

  • Core AIOps concepts and their advantages
  • The role of Prometheus and Grafana in the observability stack
  • The place of ML in AIOps: predictive versus reactive analysis

Configuring Prometheus and Grafana

  • Installing and setting up Prometheus for time series data collection
  • Building Grafana dashboards using live metrics
  • Managing exporters, relabeling, and service discovery

Preparing Data for Machine Learning

  • Extracting and processing Prometheus metrics
  • Curating datasets suitable for anomaly detection and forecasting
  • Applying Grafana transformations or Python-based data pipelines

Utilizing Machine Learning for Anomaly Detection

  • Foundational ML models for outlier identification (e.g., Isolation Forest, One-Class SVM)
  • Training and assessing models against time series data
  • Displaying detected anomalies within Grafana dashboards

Forecasting Metrics via Machine Learning

  • Developing forecasting models (ARIMA, Prophet, and an introduction to LSTM)
  • Anticipating system load and resource consumption
  • Leveraging predictions for proactive alerting and scaling strategies

Combining ML with Alerting and Automation

  • Establishing alert rules driven by ML insights or static thresholds
  • Implementing Alertmanager and notification routing
  • Initiating scripts or automation workflows upon anomaly detection

Scaling and Implementing AIOps in Production

  • Connecting external observability tools (e.g., ELK stack, Moogsoft, Dynatrace)
  • Operationalizing ML models within observability workflows
  • Best practices for scaling AIOps implementations

Recap and Future Pathways

Requirements

  • A solid grasp of system monitoring and observability fundamentals.
  • Practical experience with Grafana or Prometheus.
  • Proficiency in Python and an understanding of basic machine learning concepts.

Target Audience

  • Observability engineers.
  • Infrastructure and DevOps teams.
  • Monitoring platform architects and Site Reliability Engineers (SREs).
 14 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories