Get in Touch
 Duration 35 hours

Course Outline

Introduction, Objectives, and Migration Strategy

  • Course goals, alignment with participant profiles, and definition of success criteria.
  • High-level migration approaches and associated risk considerations.
  • Setup of workspaces, repositories, and lab datasets.

Day 1 — Migration Fundamentals and Architecture

  • Overview of Lakehouse concepts, Delta Lake, and Databricks architecture.
  • Differences between SMP and MPP and their implications for migration.
  • Medallion (Bronze→Silver→Gold) design principles and an overview of Unity Catalog.

Day 1 Lab — Translating a Stored Procedure

  • Practical migration of a sample stored procedure into a notebook environment.
  • Mapping temp tables and cursors to DataFrame transformations.
  • Validation and comparison of results against the original output.

Day 2 — Advanced Delta Lake & Incremental Loading

  • ACID transactions, commit logs, versioning, and time travel features.
  • Auto Loader, MERGE INTO patterns, upserts, and schema evolution.
  • Storage optimization techniques: OPTIMIZE, VACUUM, Z-ORDER, and partitioning.

Day 2 Lab — Incremental Ingestion & Optimization

  • Implementation of Auto Loader ingestion and MERGE workflows.
  • Application of OPTIMIZE, Z-ORDER, and VACUUM, followed by result validation.
  • Measurement of read/write performance improvements.

Day 3 — SQL in Databricks, Performance & Debugging

  • Analytical SQL features: window functions, higher-order functions, and JSON/array handling.
  • Interpreting Spark UI: DAGs, shuffles, stages, tasks, and diagnosing bottlenecks.
  • Query tuning patterns: broadcast joins, hints, caching, and spill reduction.

Day 3 Lab — SQL Refactoring & Performance Tuning

  • Refactoring a resource-intensive SQL process into optimized Spark SQL.
  • Utilizing Spark UI traces to identify and resolve skew and shuffle issues.
  • Benchmarking performance before and after tuning, and documenting steps.

Day 4 — Tactical PySpark: Replacing Procedural Logic

  • Spark execution model: driver, executors, lazy evaluation, and partitioning strategies.
  • Transforming loops and cursors into vectorized DataFrame operations.
  • Modularization, UDFs/pandas UDFs, widgets, and creating reusable libraries.

Day 4 Lab — Refactoring Procedural Scripts

  • Refactoring a procedural ETL script into modular PySpark notebooks.
  • Introducing parametrization, unit-style tests, and reusable functions.
  • Conducting code reviews and applying best-practice checklists.

Day 5 — Orchestration, End-to-End Pipeline & Best Practices

  • Databricks Workflows: job design, task dependencies, triggers, and error handling.
  • Designing incremental Medallion pipelines with quality rules and schema validation.
  • Integration with Git (GitHub/Azure DevOps), CI, and testing strategies for PySpark logic.

Day 5 Lab — Build a Complete End-to-End Pipeline

  • Assembling a Bronze→Silver→Gold pipeline orchestrated with Workflows.
  • Implementing logging, auditing, retries, and automated validations.
  • Executing the full pipeline, validating outputs, and preparing deployment notes.

Operationalization, Governance, and Production Readiness

  • Unity Catalog governance, lineage, and access control best practices.
  • Cost management, cluster sizing, autoscaling, and job concurrency patterns.
  • Deployment checklists, rollback strategies, and runbook creation.

Final Review, Knowledge Transfer, and Next Steps

  • Participant presentations on migration work and key lessons learned.
  • Gap analysis, recommended follow-up activities, and handoff of training materials.
  • References, further learning paths, and support options.

Requirements

  • A solid understanding of core data engineering concepts.
  • Hands-on experience with SQL and stored procedures (e.g., Synapse or SQL Server).
  • Familiarity with ETL orchestration frameworks (such as ADF or similar tools).

Target Audience

  • Technology managers with a data engineering background.
  • Data engineers looking to transition procedural OLAP logic to Lakehouse patterns.
  • Platform engineers responsible for overseeing Databricks adoption.

Number of participants


Price per participant

Upcoming Courses

Related Categories