Get in Touch
 Duration 21 hours

Course Outline

Introduction:

  • Apache Spark within the Hadoop Ecosystem
  • Brief overview of Python and Scala

Core Concepts (Theory):

  • Architecture
  • RDD
  • Transformations and Actions
  • Stages, Tasks, and Dependencies

Exploring Fundamentals via Databricks (Hands-on Workshop):

  • Practical exercises using the RDD API
  • Essential action and transformation functions
  • Working with PairRDD
  • Join operations
  • Caching strategies
  • Practical exercises using the DataFrame API
  • SparkSQL
  • DataFrame operations: select, filter, group, and sort
  • Implementing UDFs (User Defined Functions)
  • Exploring the Dataset API
  • Streaming data

Deployment Strategies via AWS (Hands-on Workshop):

  • Fundamentals of AWS Glue
  • Distinguishing between AWS EMR and AWS Glue
  • Implementing sample jobs in both environments
  • Evaluating advantages and limitations of each service

Additional Topics:

  • Introduction to Apache Airflow orchestration

Requirements

Programming skills (ideally in Python or Scala)

Foundational knowledge of SQL

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories