Get in Touch

Course Outline

Introduction:

  • Apache Spark within the Hadoop Ecosystem
  • Brief overview of Python and Scala

Core Concepts (Theoretical Foundations):

  • Spark Architecture
  • RDDs
  • Transformations and Actions
  • Stages, Tasks, and Dependencies

Databricks Environment: Mastering Fundamentals (Hands-on Workshop):

  • Practical exercises using the RDD API
  • Essential action and transformation functions
  • PairRDDs
  • Join operations
  • Caching strategies
  • Practical exercises using the DataFrame API
  • SparkSQL integration
  • DataFrame operations: select, filter, group, and sort
  • UDFs (User-Defined Functions)
  • Exploring the DataSet API
  • Structured Streaming

AWS Environment: Understanding Deployment (Hands-on Workshop):

  • Fundamentals of AWS Glue
  • Analysing the differences between AWS EMR and AWS Glue
  • Running example jobs in both environments
  • Evaluating advantages and limitations

Additional Topics:

  • Introduction to orchestration with Apache Airflow

Requirements

Programming proficiency (ideally in Python and Scala)

Foundational knowledge of SQL

 21 Hours

Custom Corporate Training

Training solutions designed exclusively for businesses.

  • Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
  • Flexible Schedule: Dates and times adapted to your team's agenda.
  • Format: Online (live), In-company (at your offices), or Hybrid.
Investment

Price per private group, online live training, starting from 3900 € + VAT*

Contact us for an exact quote and to hear our latest promotions

Testimonials (3)

Provisional Upcoming Courses (Contact Us For More Information)

Related Categories