DatabricksIntermediate
Databricks Lakehouse Engineering with PySpark
Build medallion-architecture pipelines on Databricks with PySpark and Delta Lake.
Syllabus · 5 modules · Weekday evenings, 7–9pm IST
- Lakehouse fundamentals · 4h
Workspace, clusters, notebooks, Delta Lake basics, medallion architecture - PySpark for data engineering · 8h
DataFrames, joins, window functions, UDF trade-offs, performance basics - Delta Lake in depth · 6h
MERGE, change data feed, schema evolution, OPTIMIZE and Z-ORDER, time travel - Governance and quality · 5h
Unity Catalog, expectations and data quality checks, lineage - Orchestration and capstone · 7h
Workflows, parameters, monitoring; capstone pipeline reviewed by the trainer
You will be able to: Build medallion (Bronze, Silver, Gold) pipelines with PySpark and Delta Lake · Handle incremental and historical loads, schema evolution and data quality checks · Govern data with Unity Catalog · Orchestrate and monitor jobs with Databricks Workflows
Fees on requestEnquire for fees and the next batch