Get in Touch

Course Outline

Fundamentals of AIOps with Open Source Tools

  • Core AIOps concepts and their organizational benefits
  • The role of Prometheus and Grafana in the observability ecosystem
  • Positioning ML within AIOps: predictive vs. reactive analytics

Deployment of Prometheus and Grafana

  • Installing and configuring Prometheus for time-series data collection
  • Building dashboards in Grafana utilizing live metrics
  • Investigating exporters, relabeling, and service discovery mechanisms

Data Preparation for Machine Learning

  • Extraction and transformation of Prometheus metrics
  • Preparing datasets suitable for anomaly detection and forecasting
  • Utilizing Grafana transformations or Python-based pipelines

Leveraging Machine Learning for Anomaly Detection

  • Fundamental ML models for outlier identification (e.g., Isolation Forest, One-Class SVM)
  • Training and assessing models on time-series datasets
  • Displaying detected anomalies within Grafana dashboards

Metric Forecasting Using ML

  • Developing simple forecasting models (intro to ARIMA, Prophet, LSTM)
  • Anticipating system load and resource consumption
  • Applying predictions for proactive alerting and scaling strategies

Integrating ML with Alerting and Automation

  • Establishing alert rules based on ML outputs or defined thresholds
  • Utilizing Alertmanager and notification routing
  • Activating scripts or automation workflows upon anomaly detection

Scaling and Implementing AIOps Operations

  • Connecting external observability tools (e.g., ELK stack, Moogsoft, Dynatrace)
  • Operationalizing ML models within observability pipelines
  • Best practices for AIOps at scale

Recap and Future Directions

Requirements

  • A solid grasp of system monitoring and observability principles
  • Practical experience with Grafana or Prometheus
  • Knowledge of Python and foundational machine learning concepts

Target Audience

  • Observability engineers
  • Infrastructure and DevOps teams
  • Monitoring platform architects and Site Reliability Engineers (SREs)
 14 Hours

Upcoming Courses

Related Categories