코스 개요
Introduction
Comprehending Big Data
Spark Overview
Python Overview
PySpark Overview
- Distributing Data Using the Resilient Distributed Datasets (RDD) Framework
- Distributing Computation Using Spark API Operators
Configuring Python with Spark
Setting Up PySpark
Utilizing Amazon Web Services (AWS) EC2 Instances for Spark
Configuring Databricks
Setting Up the AWS EMR Cluster
Mastering Fundamentals of Python Programming
- Getting Started with Python
- Utilizing the Jupyter Notebook
- Working with Variables and Basic Data Types
- Manipulating Lists
- Implementing Conditional Statements
- Handling User Inputs
- Working with while Loops
- Implementing Functions
- Working with Classes
- Managing Files and Exceptions
- Working with Projects, Data, and APIs
Exploring Spark DataFrames
- Introduction to Spark DataFrames
- Executing Basic Operations in Spark
- Applying Groupby and Aggregate Operations
- Handling Timestamps and Date Data
Spark DataFrame Project Exercise
Machine Learning with MLlib
Integrating MLlib, Spark, and Python for Machine Learning
Regression Analysis
- Theory of Linear Regression
- Implementing Regression Evaluation Code
- Linear Regression Practical Exercise
- Theory of Logistic Regression
- Implementing Logistic Regression Code
- Logistic Regression Practical Exercise
Random Forests and Decision Trees
- Theory Behind Tree-Based Methods
- Implementing Decision Trees and Random Forest Algorithms
- Random Forest Classification Practical Exercise
K-means Clustering
- Understanding K-means Clustering Theory
- Implementing K-means Clustering Code
- Clustering Practical Exercise
Recommender Systems
Implementing Natural Language Processing
- Concepts of Natural Language Processing (NLP)
- NLP Tooling Overview
- NLP Practical Exercise
Spark Streaming with Python
- Overview of Spark Streaming
- Spark Streaming Practical Exercise
요건
- Foundational programming skills
Target Audience
- Software Developers
- IT Professionals
- Data Scientists
회원 평가 (6)
실제로 적용할 수 있어 좋았습니다. 이론적인 지식을 실제 예를 통해 적용하는 것이 매우 좋아졌습니다.
Aurelia-Adriana - Allianz Services Romania
코스 - Python and Spark for Big Data (PySpark)
기계 번역됨
이 코스는 매우 복잡하고 관련된 주제들의 시리즈였으며, 파블로는 각각의 주제에 대해 깊은 전문성을 가지고 있었습니다. 때로는 미묘한 부분들이 의사소통 과정이나 시간 압박으로 인해 빠지는 경우가 있었고, 이로 인해 기대치가 완전히 충족되지 않았을 수도 있습니다. 또한 UHG/Azure Databricks 설정 문제도 있었지만, 파블로와 UHG는 문제가 발생하자마자 신속하게 해결했습니다. 이는 저에게 UHG와 파블로 간의 높은 수준의 이해와 전문성을 보여주었습니다,
Michael Monks - Tech NorthWest Skillnet
코스 - Python and Spark for Big Data (PySpark)
기계 번역됨
개별적인 주의.
ARCHANA ANILKUMAR - PPL
코스 - Python and Spark for Big Data (PySpark)
기계 번역됨
실습 훈련..
Abraham Thomas - PPL
코스 - Python and Spark for Big Data (PySpark)
기계 번역됨
수업은 Jupyter 노트북에서 진행되었습니다. 주제들은 논리적인 순서로 구성되어 있으며, 덕분에 간단한 부분부터 복잡한 부분까지 자연스럽게 세션을 발전시킬 수 있었습니다. 저는 이미 머신 러닝 배경을 가진 고급 파이썬 사용자여서, 이 과정을 possibly 몇몇 동료 학생들이 훈련 과정을 받았던 것보다 더 쉽게 따라갈 수 있었습니다. 가장 기초적인 개념들을 생략하고 가장 중요한 사항에 집중한 점을 높이 평가합니다.
Angela DeLaMora - ADT, LLC
코스 - Python and Spark for Big Data (PySpark)
기계 번역됨
연습 과제
Pawel Kozikowski - GE Medical Systems Polska Sp. Zoo
코스 - Python and Spark for Big Data (PySpark)
기계 번역됨