Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps with Open Source Tools
- Exploration of AIOps core concepts and their operational benefits
- The role of Prometheus and Grafana within the observability stack
- Positioning ML in AIOps: distinguishing between predictive and reactive analytics
Setting Up Prometheus and Grafana
- Installation and configuration of Prometheus for effective time series collection
- Designing Grafana dashboards that utilize real-time metrics
- In-depth review of exporters, relabeling, and service discovery mechanisms
Data Preprocessing for ML
- Techniques for extracting and transforming Prometheus metrics
- Preparing datasets optimized for anomaly detection and forecasting tasks
- Utilizing Grafana’s native transformations or Python-based pipelines
Applying Machine Learning for Anomaly Detection
- Introduction to basic ML models for outlier detection, such as Isolation Forest and One-Class SVM
- Processes for training and evaluating models against time series data
- Strategies for visualizing detected anomalies within Grafana dashboards
Forecasting Metrics with ML
- Development of simple forecasting models, including an introduction to ARIMA, Prophet, and LSTM
- Methods for predicting system load and resource usage patterns
- Leveraging predictions to drive early alerting and scaling decisions
Integrating ML with Alerting and Automation
- Formulating alert rules based on ML outputs or dynamic thresholds
- Implementing Alertmanager and configuring notification routing
- Automating workflows and script triggers in response to detected anomalies
Scaling and Operationalizing AIOps
- Integration with external observability tools, such as the ELK stack, Moogsoft, and Dynatrace
- Strategies for operationalizing ML models within observability pipelines
- Best practices for implementing AIOps at scale
Summary and Next Steps
Requirements
- Solid understanding of system monitoring and observability principles
- Practical experience with Grafana or Prometheus
- Working knowledge of Python and foundational machine learning concepts
Target Audience
- Observability engineers
- Infrastructure and DevOps teams
- Monitoring platform architects and site reliability engineers (SREs)