Data Private

Introduction to Spark with Python (SPK103)

3 days

Introduction to Spark with Python provides a practical foundation in Apache Spark for learners who want to process, analyze, and work with large-scale data using Spark’s distributed computing capabilities.

Register or Request Training

  • Private class for your team
  • Live expert instructor
  • Online or on‑location
  • Customizable agenda
  • Proposal responses same day as request

Course Overview

Introduction to Spark with Python provides a practical foundation in Apache Spark for learners who want to process, analyze, and work with large-scale data using Spark’s distributed computing capabilities. The course begins with the Spark ecosystem, architecture, installation, SparkContext, and the Spark shell before exploring resilient distributed datasets (RDDs), lazy evaluation, partitioning, transformations, and common data-processing operations.

Learners then work with Spark SQL and DataFrames, including loading and saving common data formats, schema inference, filtering, grouping, aggregation, SQL queries, and data transformations. The course also examines shuffling, query optimization with Catalyst and Tungsten, caching, broadcast variables, performance tuning, and Spark application architecture. By the end of the course, learners will also understand how standalone Spark applications are configured and executed and how Structured Streaming can process continuous data, including streams sourced from Kafka.

Course Benefits

  • Explore the Spark ecosystem, architecture, core components, and key differences between Spark and Hadoop.
  • Set up Spark and work with the Spark shell, SparkContext, and SparkSession.
  • Create and transform RDDs using operations such as map and filter while understanding partitioning and lazy evaluation.
  • Load, save, query, and transform structured data with Spark SQL and DataFrames.
  • Use common data formats including JSON, CSV, text, and Parquet in Spark applications.
  • Analyze grouping, reducing, joining, shuffling, and dependency patterns to understand their impact on application performance.
  • Apply Catalyst, Tungsten, caching, broadcast variables, and other techniques to improve Spark processing efficiency.
  • Configure and build standalone Spark applications and understand the roles of drivers, executors, tasks, and cluster managers.
  • Practice diagnosing Spark applications through logging, debugging, query plans, and general performance-tuning techniques.
  • Develop Structured Streaming workflows that consume, process, and output continuous data, including data from Kafka.

Delivery Methods

Private Class
Delivered for your team at your site or online.

Course Outline

  1. Introduction to Spark
    1. Overview, Motivations, Spark Systems
    2. Spark Ecosystem
    3. Spark vs. Hadoop
    4. Acquiring and Installing Spark
    5. The Spark Shell, SparkContext
  2. RDDs and Spark Architecture
    1. RDD Concepts, Lifecycle, Lazy Evaluation
    2. RDD Partitioning and Transformations
    3. Working with RDDs - Creating and Transforming (map, filter, etc.)
  3. Spark SQL, DataFrames, and DataSets
    1. Overview
    2. SparkSession, Loading/Saving Data, Data Formats (JSON, CSV, Parquet, text ...)
    3. Introducing DataFrames (Creation and Schema Inference)
    4. Supported Data Formats (JSON, Text, CSV, Parquet)
    5. Working with the DataFrame (untyped) Query DSL (Column, Filtering, Grouping, Aggregation)
    6. SQL-based Queries
    7. Mapping and Splitting (flatMap(), explode(), and split())
    8. DataFrames vs. RDDs
  4. Shuffling Transformations and Performance
    1. Grouping, Reducing, Joining
    2. Shuffling, Narrow vs. Wide Dependencies, and Performance Implications
    3. Exploring the Catalyst Query Optimizer (explain(), Query Plans, Issues with lambdas)
    4. The Tungsten Optimizer (Binary Format, Cache Awareness, Whole-Stage Code Gen)
  5. Performance Tuning
    1. Caching - Concepts, Storage Type, Guidelines
    2. Minimizing Shuffling for Increased Performance
    3. Using Broadcast Variables and Accumulators
    4. General Performance Guidelines
  6. Creating Standalone Applications
    1. Core API, SparkSession.Builder
    2. Configuring and Creating a SparkSession
    3. Building and Running Applications - sbt/build.sbt and spark-submit
    4. Application Lifecycle (Driver, Executors, and Tasks)
    5. Cluster Managers (Standalone, YARN, Mesos)
    6. Logging and Debugging
  7. Spark Streaming
    1. Introduction and Streaming Basics
    2. Streaming Introduction
    3. Structured Streaming (Spark 2+)
    4. Continuous Applications
    5. Table Paradigm, Result Table
    6. Steps for Structured Streaming
    7. Sources and Sinks
    8. Consuming Kafka Data
    9. Kafka Overview
    10. Structured Streaming - "kafka" format
    11. Processing the Stream

Class Materials

Each student receives a comprehensive set of materials, including course notes and all class examples.

Class Prerequisites

Experience in the following is required for this Spark class:

  • Working knowledge of some programming language. No Python experience necessary.

Have questions about this course?

We can help with curriculum details, delivery options, pricing, or anything else. Reach out and we’ll point you in the right direction.