Over the course of three intensive days, this practical training programme centres on the development and refinement of high-efficiency data processing workflows leveraging PySpark, Pandas, and Polars within Kubernetes-based ecosystems.
Learners will cultivate a hands-on grasp of Spark application execution on Kubernetes, specifically examining how configuration decisions at the application level impact performance, scalability, resource utilization, and operational costs. Key optimisation topics include sizing executors, managing memory allocation, dynamic allocation, partitioning methodologies, shuffle dynamics, mitigating small-file issues, and enhancing Parquet processing efficiency.
Additionally, the course tackles frequent obstacles encountered in Pandas usage, such as memory constraints and out-of-memory exceptions, while introducing Polars as a high-performance alternative for specific data processing tasks. Through practical exercises, participants will learn to identify performance and memory bottlenecks, evaluate various configuration strategies, and implement optimisation techniques in realistic ETL and machine learning contexts.
The primary focus remains on practical decision-making: gaining the ability to pinpoint bottlenecks, select the most suitable tool, configure Spark for peak efficiency, and strike a balance between performance and infrastructure resource consumption and cost.
Read more...