Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 35 hours
Course Outline
Introduction, Learning Objectives, and Migration Strategy
- Course objectives, alignment with participant profiles, and key success metrics
- Overview of high-level migration approaches and associated risk factors
- Configuration of workspaces, repositories, and lab datasets
Day 1 — Migration Fundamentals and Architecture
- Core Lakehouse concepts, Delta Lake introduction, and Databricks architecture
- Distinguishing SMP vs MPP models and their impact on migration strategies
- Medallion (Bronze→Silver→Gold) architecture design and Unity Catalog overview
Day 1 Lab — Translating a Stored Procedure
- Practical migration of a sample stored procedure into a notebook
- Converting temp tables and cursors into DataFrame transformations
- Validation and comparison against original output results
Day 2 — Advanced Delta Lake & Incremental Loading
- ACID transactions, commit logs, versioning, and time travel features
- Auto Loader, MERGE INTO patterns, upserts, and schema evolution
- Optimization techniques: OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage tuning
Day 2 Lab — Incremental Ingestion & Optimization
- Implementing Auto Loader ingestion and MERGE workflows
- Applying OPTIMIZE, Z-ORDER, and VACUUM commands; verifying results
- Assessing read/write performance enhancements
Day 3 — SQL in Databricks, Performance & Debugging
- Analytical SQL capabilities: window functions, higher-order functions, and JSON/array handling
- Interpreting Spark UI: DAGs, shuffles, stages, tasks, and bottleneck identification
- Query tuning strategies: broadcast joins, hints, caching, and reducing spill
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring complex SQL processes into optimized Spark SQL
- Using Spark UI traces to diagnose and resolve skew and shuffle issues
- Benchmarking pre- and post-optimization performance and documenting steps
Day 4 — Tactical PySpark: Replacing Procedural Logic
- Spark execution model: driver, executors, lazy evaluation, and partitioning strategies
- Transforming loops and cursors into vectorized DataFrame operations
- Modularization techniques, UDFs/pandas UDFs, widgets, and creating reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Converting procedural ETL scripts into modular PySpark notebooks
- Incorporating parametrization, unit-style tests, and reusable functions
- Conducting code reviews and applying best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Databricks Workflows: job design, task dependencies, triggers, and error handling
- Designing incremental Medallion pipelines with quality rules and schema validation
- Integrating with Git (GitHub/Azure DevOps), CI, and testing strategies for PySpark logic
Day 5 Lab — Building a Complete End-to-End Pipeline
- Assembling a Bronze→Silver→Gold pipeline orchestrated via Workflows
- Implementing logging, auditing, retries, and automated validations
- Executing the full pipeline, validating outputs, and preparing deployment documentation
Operationalization, Governance, and Production Readiness
- Best practices for Unity Catalog governance, lineage, and access controls
- Cost management, cluster sizing, autoscaling, and job concurrency patterns
- Creating deployment checklists, rollback strategies, and operational runbooks
Final Review, Knowledge Transfer, and Next Steps
- Participant presentations on migration work and key takeaways
- Gap analysis, recommended follow-up actions, and handover of training materials
- References, advanced learning paths, and support options
Requirements
- A solid grasp of data engineering concepts
- Practical experience with SQL and stored procedures (Synapse / SQL Server)
- Knowledge of ETL orchestration concepts (ADF or similar tools)
Target Audience
- Technology managers with a background in data engineering
- Data engineers transitioning from procedural OLAP logic to Lakehouse patterns
- Platform engineers overseeing Databricks adoption