Get in Touch
 Duration 14 hours

Course Outline

Introduction to AIOps

  • Defining AIOps and its strategic value.
  • Comparing traditional monitoring with AIOps-driven observability.
  • Examining AIOps architecture and essential components.

Operational Data Collection and Normalization

  • Identifying observability data types: metrics, logs, and traces.
  • Ingesting data from diverse sources such as servers, containers, and cloud environments.
  • Deploying agents and exporters (e.g., Prometheus, Beats, Fluentd).

Data Correlation and Anomaly Detection

  • Applying time series correlation and statistical analysis.
  • Leveraging ML models for anomaly identification.
  • Detecting incidents within distributed system landscapes.

Alert Management and Noise Mitigation

  • Crafting intelligent alert rules and dynamic thresholds.
  • Implementing suppression, deduplication, and alert grouping strategies.
  • Integrating with platforms like Alertmanager, Slack, PagerDuty, or Opsgenie.

Root Cause Analysis and Visualization

  • Utilizing dashboards to visualize metrics and uncover trends.
  • Analyzing event timelines to support RCA.
  • Tracing issues across layers using distributed tracing tools.

Automation and Remediation Strategies

  • Triggering automated scripts or workflows in response to incidents.
  • Integrating with ITSM systems such as ServiceNow or Jira.
  • Exploring use cases including self-healing, scaling, and traffic rerouting.

Open Source and Commercial AIOps Ecosystems

  • Surveying key tools: Prometheus, Grafana, ELK, Moogsoft, and Dynatrace.
  • Evaluating criteria for selecting an appropriate AIOps platform.
  • Conducting demonstrations and hands-on sessions with a chosen stack.

Summary and Path Forward

Requirements

  • A solid foundation in IT operations and system monitoring concepts.
  • Hands-on experience with monitoring tools or dashboard interfaces.
  • Basic familiarity with standard log and metric formats.

Target Audience

  • Operations teams managing infrastructure and application stacks.
  • Site Reliability Engineers (SREs).
  • Teams dedicated to IT monitoring and observability.

Number of participants


Price per participant

Upcoming Courses

Related Categories