Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to AIOps
- Defining AIOps and its significance.
- Comparing traditional monitoring with AIOps-driven observability.
- Exploring AIOps architecture and key components.
Collecting and Normalizing Operational Data
- Types of observability data: metrics, logs, and traces.
- Ingesting data from various sources (servers, containers, cloud).
- Implementing agents and exporters (Prometheus, Beats, Fluentd).
Data Correlation and Anomaly Detection
- Time series correlation and statistical methodologies.
- Applying ML models for anomaly detection.
- Identifying incidents within distributed systems.
Alerting and Noise Reduction
- Designing intelligent alert rules and thresholds.
- Implementing suppression, deduplication, and alert grouping.
- Integrations with Alertmanager, Slack, PagerDuty, or Opsgenie.
Root Cause Analysis and Visualization
- Utilizing dashboards to visualize metrics and identify trends.
- Analyzing events and timelines for RCA.
- Tracking issues across layers using distributed tracing tools.
Automation and Remediation
- Initiating automated scripts or workflows triggered by incidents.
- Connecting with ITSM systems (ServiceNow, Jira).
- Application examples: self-healing, scaling, and traffic rerouting.
Open Source and Commercial AIOps Platforms
- Overview of tools: Prometheus, Grafana, ELK, Moogsoft, Dynatrace.
- Criteria for evaluating and selecting an AIOps platform.
- Demonstration and hands-on practice with a chosen stack.
Summary and Next Steps
Requirements
- A solid understanding of IT operations and system monitoring concepts.
- Prior experience with monitoring tools or dashboards.
- Familiarity with basic log and metric formats.
Audience
- Operations teams managing infrastructure and applications.
- Site Reliability Engineers (SREs).
- IT monitoring and observability teams.
14 Hours