This intensive three-day workshop is designed to help you build and fine-tune high-efficiency data processing workflows utilizing PySpark, Pandas, and Polars within Kubernetes-based ecosystems.
Attendees will gain a deep, hands-on grasp of how Spark applications function on Kubernetes. You will learn how configuration choices at the application level directly impact performance, scalability, resource utilization, and operational costs. The curriculum delves into critical optimization strategies, such as determining optimal executor sizing, managing memory allocation, leveraging dynamic allocation, refining partitioning tactics, understanding shuffle dynamics, resolving the small-file challenge, and achieving efficient Parquet processing.
The training also tackles frequent hurdles encountered when using Pandas, specifically memory constraints and out-of-memory errors. It introduces Polars as a high-performance alternative for specific data processing tasks. Through practical exercises, participants will learn to diagnose performance and memory bottlenecks, evaluate various configuration approaches, and apply optimization techniques to real-world ETL and machine learning use cases.
Throughout the course, the primary focus remains on practical decision-making: mastering the ability to pinpoint bottlenecks, choose the right tool, configure Spark for maximum efficiency, and strike a balance between high performance and the consumption of infrastructure resources and costs.
Read more...