Get in Touch
 Duration 14 hours

Course Outline

Introduction to AIOps with Open Source Tools

  • Exploration of AIOps core concepts and their operational benefits
  • The role of Prometheus and Grafana within the observability stack
  • Positioning ML in AIOps: distinguishing between predictive and reactive analytics

Setting Up Prometheus and Grafana

  • Installation and configuration of Prometheus for effective time series collection
  • Designing Grafana dashboards that utilize real-time metrics
  • In-depth review of exporters, relabeling, and service discovery mechanisms

Data Preprocessing for ML

  • Techniques for extracting and transforming Prometheus metrics
  • Preparing datasets optimized for anomaly detection and forecasting tasks
  • Utilizing Grafana’s native transformations or Python-based pipelines

Applying Machine Learning for Anomaly Detection

  • Introduction to basic ML models for outlier detection, such as Isolation Forest and One-Class SVM
  • Processes for training and evaluating models against time series data
  • Strategies for visualizing detected anomalies within Grafana dashboards

Forecasting Metrics with ML

  • Development of simple forecasting models, including an introduction to ARIMA, Prophet, and LSTM
  • Methods for predicting system load and resource usage patterns
  • Leveraging predictions to drive early alerting and scaling decisions

Integrating ML with Alerting and Automation

  • Formulating alert rules based on ML outputs or dynamic thresholds
  • Implementing Alertmanager and configuring notification routing
  • Automating workflows and script triggers in response to detected anomalies

Scaling and Operationalizing AIOps

  • Integration with external observability tools, such as the ELK stack, Moogsoft, and Dynatrace
  • Strategies for operationalizing ML models within observability pipelines
  • Best practices for implementing AIOps at scale

Summary and Next Steps

Requirements

  • Solid understanding of system monitoring and observability principles
  • Practical experience with Grafana or Prometheus
  • Working knowledge of Python and foundational machine learning concepts

Target Audience

  • Observability engineers
  • Infrastructure and DevOps teams
  • Monitoring platform architects and site reliability engineers (SREs)

Number of participants


Price per participant

Upcoming Courses

Related Categories