Cloud Private Public

Implement data engineering solutions using Azure Databricks (DP-750T00)

4 days

Build, secure, govern, deploy, and optimize scalable lakehouse solutions using Azure Databricks, Unity Catalog, Lakeflow, and Apache Spark for data engineering.

Register or Request Training

Price per student
$2,445.10
Guaranteed to run
Select a date
Please select a class.
  • Private class for your team
  • Live expert instructor
  • Online or on‑location
  • Customizable agenda
  • Proposal responses same day as request

Course Overview

Master end-to-end data engineering with Azure Databricks and Unity Catalog. Learn to configure environments, build robust ingestion and transformation pipelines, implement enterprise governance and security, and deploy optimized workloads. By the end of this course, you will be prepared to implement, secure, monitor, and maintain scalable lakehouse solutions.

Course Benefits

  • Configure Azure Databricks compute, storage, and integrations for data engineering workloads.
  • Create, organize, secure, and govern data assets with Unity Catalog.
  • Design data models, partitioning schemes, clustering strategies, and slowly changing dimensions.
  • Ingest batch and streaming data using Lakeflow, notebooks, SQL, Auto Loader, and Spark Structured Streaming.
  • Cleanse, transform, and load data while managing schema drift and data quality constraints.
  • Design, build, schedule, and manage Azure Databricks pipelines and Lakeflow Jobs.
  • Apply Git-based development, testing, packaging, and deployment practices.
  • Monitor, troubleshoot, and optimize Azure Databricks workloads.

Delivery Methods

Public Class
Live expert-led online training from anywhere. Guaranteed to run .
Private Class
Delivered for your team at your site or online.

Course Outline

  1. Explore Azure Databricks
    1. Get started with Azure Databricks
    2. Identify Azure Databricks workloads
    3. Understand key concepts
    4. Explore data governance using Unity Catalog and Microsoft Purview
    5. Module assessment
  2. Understand Azure Databricks Architecture
    1. Understand Azure Databricks architecture
    2. Understand Unity Catalog managed storage
    3. Understand external storage
    4. Understand default storage
    5. Module assessment
  3. Understand Azure Databricks Integrations
    1. Understand integration with Microsoft Fabric
    2. Understand integration with Power BI
    3. Understand integration with VS Code
    4. Understand integration with Power Platform
    5. Understand integration with Copilot Studio
    6. Understand integration with Microsoft Purview
    7. Understand integration with Microsoft Foundry
    8. Module assessment
  4. Select and Configure Compute in Azure Databricks
    1. Choose an appropriate compute type
    2. Configure compute performance
    3. Configure compute features
    4. Install libraries for compute
    5. Configure compute access
    6. Module assessment
  5. Create and Organize Objects in Unity Catalog
    1. Apply naming conventions
    2. Create a catalog
    3. Create a schema
    4. Create tables and views
    5. Create volumes
    6. Implement DDL operations
    7. Implement a foreign catalog
    8. Configure AI/BI Genie instructions
  6. Secure Unity Catalog Objects
    1. Understand the query lifecycle
    2. Implement access control strategies
    3. Understand fine-grained access control
    4. Implement row filtering and column masking
    5. Access Azure Key Vault secrets
    6. Authenticate data access with service principals
    7. Authenticate resource access with managed identities
    8. Module assessment
  7. Govern Unity Catalog Objects
    1. Create and preserve table definitions
    2. Configure attribute-based access control with tags and policies
    3. Apply data retention policies
    4. Set up and manage data lineage
    5. Configure audit logging
    6. Design a secure Delta Sharing strategy
    7. Module assessment
  8. Design and Implement Data Modeling with Azure Databricks
    1. Design ingestion logic and data source configuration
    2. Choose a data ingestion tool
    3. Choose a data table format
    4. Design and implement a data partitioning scheme
    5. Choose a slowly changing dimension type
    6. Implement a type 2 slowly changing dimension
    7. Design and implement a temporal table to record changes over time
    8. Choose column or table granularity based on requirements
    9. Choose between managed and external tables
    10. Design and implement a clustering strategy
  9. Ingest Data into Unity Catalog
    1. Ingest data with Lakeflow Connect
    2. Ingest data with notebooks
    3. Ingest data with SQL methods
    4. Ingest data with a change data capture feed
    5. Ingest data with Spark Structured Streaming
    6. Ingest data with Auto Loader
    7. Ingest data with Lakeflow Spark Declarative Pipelines
    8. Module assessment
  10. Cleanse, Transform, and Load Data into Unity Catalog
    1. Profile data
    2. Choose column data types
    3. Resolve duplicates and nulls
    4. Transform data with filters and aggregations
    5. Transform data with joins and set operators
    6. Transform data with denormalization and pivots
    7. Load data with merge, insert, and append operations
    8. Module assessment
  11. Implement and Manage Data Quality Constraints with Azure Databricks
    1. Implement validation checks
    2. Implement data type checks
    3. Detect and manage schema drift
    4. Manage data quality with pipeline expectations
    5. Module assessment
  12. Design and Implement Data Pipelines with Azure Databricks
    1. Design the order of operations for a pipeline
    2. Choose between notebooks and Lakeflow Pipelines
    3. Design Lakeflow job logic
    4. Design error handling for pipelines and jobs
    5. Create a pipeline with a notebook
    6. Create a pipeline with Lakeflow Spark Declarative Pipelines
    7. Module assessment
  13. Implement Lakeflow Jobs with Azure Databricks
    1. Create job setup and configuration
    2. Configure job triggers
    3. Schedule a job
    4. Configure job alerts
    5. Configure automatic restarts
    6. Module assessment
  14. Implement Development Lifecycle Processes in Azure Databricks
    1. Apply Git version control best practices
    2. Manage branching and pull requests
    3. Implement a testing strategy
    4. Configure and package Declarative Automation Bundles
    5. Deploy a bundle with the Databricks CLI
    6. Module assessment
  15. Monitor, Troubleshoot, and Optimize Workloads in Azure Databricks
    1. Monitor and manage cluster consumption
    2. Troubleshoot and repair Lakeflow Jobs
    3. Troubleshoot Spark jobs and notebooks
    4. Investigate caching, skewing, spilling, and shuffling
    5. Implement log streaming with Azure Log Analytics
    6. Module assessment

Class Materials

Each student receives a comprehensive set of materials, including course notes and all class examples.

Class Prerequisites

Experience in the following is required for this Azure class:

  • Ability to work with SQL.
  • Experience using Python and notebooks for data engineering tasks.
  • Understanding of Azure Databricks workspaces and Unity Catalog.
  • Familiarity with data access patterns and core data engineering and data warehouse concepts.

Experience in the following would be useful for this Azure class:

  • Fundamental knowledge of data analytics concepts.
  • Basic understanding of cloud storage and data organization principles.
  • Foundational knowledge of Azure security, including Microsoft Entra ID.
  • Familiarity with Git version control fundamentals.

Have questions about this course?

We can help with curriculum details, delivery options, pricing, or anything else. Reach out and we’ll point you in the right direction.