Hadoop Data Analyst Training

Customized Onsite Training

3
Days
  • Customized Content
  • For Groups of 5+
  • Online or On-location
  • Expert Instructors
Overview

In this Hadoop Data Analyst training class, students learn the fundamentals of Hadoop and move on to working with Pig and Hive.

Goals
  1. Understand Hadoop fundamentals.
  2. Learn to analyze data with Pig.
  3. Learn to process complex data with Pig.
  4. Learn to troubleshoot Pig.
  5. Learn when to use Hive.
  6. Learn to manage data with Hive.
  7. Learn to optimize Hive.
Outline
  1. Hadoop Fundamentals
    1. Hadoop Overview
    2. HDFS
    3. MapReduce
    4. The Hadoop Ecosystem
  2. Introduction to Pig
    1. Pig's Features/Use Cases
    2. Interacting with Pig
  3. Basic Data Analysis with Pig
    1. Pig Latin
    2. Loading Data
    3. Field Definitions and Simple Data Types
    4. Data Output
    5. Viewing the Schema
    6. Filtering /Sorting Data
    7. Common Functions
  4. Processing Complex Data with Pig
    1. Storage Formats
    2. Complex/Nested Data Types
    3. Grouping
    4. Built-in Functions for working with Complex Data
    5. Iterating Grouped Data
  5. Multi-Dataset Operations with Pig
    1. Combining Data Sets
    2. Joining Data Sets
    3. Set Operations
    4. Splitting Data Sets
  6. Extending Pig
    1. Parameters
    2. Macros / Imports
    3. UDFs
    4. Using Other Languages to Process Data with Pig
  7. Pig Troubleshooting and Optimization
    1. Logging
    2. Hadoop's Web UI
    3. Data Sampling and Debugging
    4. Understanding the Execution Plan Improving the Performance
  8. Introduction to Hive
    1. Hive Schema and Data Storage
    2. Hive vs Traditional Databases
    3. Hive vs. Pig
    4. When to use Hive
    5. Relational Data Analysis with Hive
    6. Hive Databases and Tables
    7. Basic HiveQL Syntax
    8. Data Types
    9. Joining Data Sets
    10. Common Built-in Functions
  9. Hive Data Management
    1. Hive Data Formats
    2. Creating Databases and Hive-Managed Tables
    3. Loading Data into Hive
    4. Altering Databases and Tables Self-Managed Tables
    5. Simplifying Queries with Views
    6. Storing Query Results
    7. Controlling Access to Data
  10. Text Processing with Hive
    1. Text Processing
    2. Important String Functions
    3. Using Regular Expressions in Hive
  11. Hive Optimization
    1. Understanding Query Performance
    2. Controlling Job Execution Plan
    3. Partitioning
    4. Bucketing
    5. Indexing Data
  12. Extending Hive
    1. Data Transformation with Custom Scripts
    2. User-Defined Functions
    3. Parameterized Queries
Class Materials

Each student in our Live Online and our Onsite classes receives a comprehensive set of materials, including course notes and all the class examples.

Preparing for Class

No cancelation for low enrollment

Certified Microsoft Partner

Registered Education Provider (R.E.P.)

GSA schedule pricing

80,130

Students who have taken Live Online Training

15,540

Organizations who trust Webucator for their training needs

100%

Satisfaction guarantee and retake option

9.38

Students rated our trainers 9.38 out of 10 based on 4,997 reviews

This was an excellent class. The materials were very pertinent and well organized, and the instructor was extremely knowledgeable.

Alison Wills, Protective Life
Cropwell AL

A fantastic class! Everything I learned will save time, money and ease frustration!

Allison Barns, NOAA
Seattle WA

Covered many topics that I will be using in the future. The training was hands-on with the ability to show the instructor where I was and what I was doing, so that I would be able to stay on course.

Lila Cleveland, UMass Amherst
Amherst MA

I feel like I know GA inside and out now. I am excited to bring my knowledge back and start applying within my organization.

Kate Ludwig, None
Adams MA

Contact Us or call 1-877-932-8228