Databricks · Intermediate

Databricks Lakehouse Engineering with PySpark

Build medallion-architecture pipelines on Databricks with PySpark and Delta Lake.

Learn to design and build Bronze, Silver and Gold layers on Databricks, with Delta Lake, Unity Catalog and workflow orchestration, using realistic retail and banking datasets.

Duration
30 hours
Mode
Live online
Next batch
To be announced
Schedule
Weekday evenings, 7–9pm IST

Who it is for

Data engineers and analytics engineers adopting a Lakehouse platform.

Prerequisites

Working knowledge of SQL and basic Python.

What you will be able to do

  • Build medallion (Bronze, Silver, Gold) pipelines with PySpark and Delta Lake
  • Handle incremental and historical loads, schema evolution and data quality checks
  • Govern data with Unity Catalog
  • Orchestrate and monitor jobs with Databricks Workflows

Course contents · 5 modules · 30 hours

01
Lakehouse fundamentals

Workspace, clusters, notebooks, Delta Lake basics, medallion architecture

4 h
02
PySpark for data engineering

DataFrames, joins, window functions, UDF trade-offs, performance basics

8 h
03
Delta Lake in depth

MERGE, change data feed, schema evolution, OPTIMIZE and Z-ORDER, time travel

6 h
04
Governance and quality

Unity Catalog, expectations and data quality checks, lineage

5 h
05
Orchestration and capstone

Workflows, parameters, monitoring; capstone pipeline reviewed by the trainer

7 h

Hands-on approach

30% concepts, 70% hands-on labs on realistic business scenarios. AI practice support between sessions; every project is reviewed by the trainer.