Get in Touch

Course Outline

Introduction:

  • Apache Spark within the Hadoop Ecosystem
  • Brief overview of Python and Scala

Foundational Theory:

  • Spark Architecture
  • RDD Concepts
  • Transformations and Actions
  • Stages, Tasks, and Dependencies

Databricks Environment Workshop (Hands-on):

  • Practical exercises using the RDD API
  • Implementing basic action and transformation functions
  • Working with PairRDD
  • Executing Joins
  • Optimizing with Caching Strategies
  • Practical exercises using the DataFrame API
  • Utilizing SparkSQL
  • DataFrame operations: select, filter, group, and sort
  • Creating User Defined Functions (UDFs)
  • Exploring the DataSet API
  • Introduction to Streaming

AWS Deployment Workshop (Hands-on):

  • Foundations of AWS Glue
  • Comparing AWS EMR and AWS Glue
  • Developing example jobs in both environments
  • Evaluating advantages and limitations

Additional Topics:

  • Introduction to Apache Airflow Orchestration

Requirements

Programming proficiency (Python and Scala preferred)

Basic knowledge of SQL

 21 Hours

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories