Hadoop Data Analyst Training

In this Hadoop Data Analyst training class, students learn the fundamentals of Hadoop and move on to working with Pig and Hive.

Goals
  1. Understand Hadoop fundamentals.
  2. Learn to analyze data with Pig.
  3. Learn to process complex data with Pig.
  4. Learn to troubleshoot Pig.
  5. Learn when to use Hive.
  6. Learn to manage data with Hive.
  7. Learn to optimize Hive.
Outline
  1. Hadoop Fundamentals
    1. Hadoop Overview
    2. HDFS
    3. MapReduce
    4. The Hadoop Ecosystem
  2. Introduction to Pig
    1. Pig's Features/Use Cases
    2. Interacting with Pig
  3. Basic Data Analysis with Pig
    1. Pig Latin
    2. Loading Data
    3. Field Definitions and Simple Data Types
    4. Data Output
    5. Viewing the Schema
    6. Filtering /Sorting Data
    7. Common Functions
  4. Processing Complex Data with Pig
    1. Storage Formats
    2. Complex/Nested Data Types
    3. Grouping
    4. Built-in Functions for working with Complex Data
    5. Iterating Grouped Data
  5. Multi-Dataset Operations with Pig
    1. Combining Data Sets
    2. Joining Data Sets
    3. Set Operations
    4. Splitting Data Sets
  6. Extending Pig
    1. Parameters
    2. Macros / Imports
    3. UDFs
    4. Using Other Languages to Process Data with Pig
  7. Pig Troubleshooting and Optimization
    1. Logging
    2. Hadoop's Web UI
    3. Data Sampling and Debugging
    4. Understanding the Execution Plan Improving the Performance
  8. Introduction to Hive
    1. Hive Schema and Data Storage
    2. Hive vs Traditional Databases
    3. Hive vs. Pig
    4. When to use Hive
    5. Relational Data Analysis with Hive
    6. Hive Databases and Tables
    7. Basic HiveQL Syntax
    8. Data Types
    9. Joining Data Sets
    10. Common Built-in Functions
  9. Hive Data Management
    1. Hive Data Formats
    2. Creating Databases and Hive-Managed Tables
    3. Loading Data into Hive
    4. Altering Databases and Tables Self-Managed Tables
    5. Simplifying Queries with Views
    6. Storing Query Results
    7. Controlling Access to Data
  10. Text Processing with Hive
    1. Text Processing
    2. Important String Functions
    3. Using Regular Expressions in Hive
  11. Hive Optimization
    1. Understanding Query Performance
    2. Controlling Job Execution Plan
    3. Partitioning
    4. Bucketing
    5. Indexing Data
  12. Extending Hive
    1. Data Transformation with Custom Scripts
    2. User-Defined Functions
    3. Parameterized Queries
Class Materials

Each student in our Live Online and our Onsite classes receives a comprehensive set of materials, including course notes and all the class examples.

Preparing for Class

Training for your Team

Length: 3 Days
  • Private Class for your Team
  • Online or On-location
  • Customizable
  • Expert Instructors

What people say about our training

Course was quite educational with a lot of detailed information about SSRS application.
Shefali Pravasi
Bonfils Blood Center
Our instucture was very professional, knowledgable, and versatile.
Darrin DeVries
Bronson Healtcare Group
This is a really good course to start your Oracle career.
Guojian Cai
TD Bank Canada
This class increased my overall understanding of SQL a tremendous amount.
William Blackburn
Edgewood Chemical Biological Center

No cancelation for low enrollment

Certified Microsoft Partner

Registered Education Provider (R.E.P.)

GSA schedule pricing

61,490

Students who have taken Instructor-led Training

11,769

Organizations who trust Webucator for their Instructor-led training needs

100%

Satisfaction guarantee and retake option

9.29

Students rated our trainers 9.29 out of 10 based on 29,331 reviews

Contact Us or call 1-877-932-8228