Prepare For A Data Engineering Career
Data Engineers provide a critical bridge between data generation and data utilization by creating robust data pipelines that extract information from multiple sources, cleaning and structuring the data, and making it accessible for data analysts, data scientists, and machine learning models to derive insights and support decision-making. Their work involves a combination of software engineering, database design, cloud computing, and data management skills to create scalable and efficient systems that can handle massive amounts of information for both traditional business intelligence and modern AI applications.
The Data Engineering Bootcamp offers two starting points to prepare you for these in-demand jobs. Data Engineering with Python Foundations is for beginners without programming experience. Data Engineering for Experienced Programming Professionals is for software developers, data analysts, and other working tech professionals with programming experience.
Program Highlights
- Get hands-on experience applying the data engineering lifecycle to make raw data a valuable and usable resource for organizations.
- Learn to write Python and SQL to ingest, transform, and deliver data across a variety of workflows and tools.
- Build data pipelines using real-world datasets and tackle challenges from industries like healthcare, finance, entertainment, and more.
- Work with modern tooling—Docker, Airflow, dbt, Databricks, Snowflake—and explore solutions across all major cloud platforms, including AWS, Microsoft Azure, and Google Cloud.
- Learn valuable skills in prompting within conversational AI platforms and coding agents, and how to use generative AI safely and effectively.
- Prepare for your job search with resume workshops, interview practice, Demo Day, and post-graduation career support.
Attend our virtual information session on the 3rd Wednesday of every month at 5:30 p.m.
What You Will Learn
-
Python Fundamentals
Master core Python programming required to write clean, executable code from scratch. You will learn syntax, data types, collections, control flow, modular functions, exceptions, and file handling across common formats (CSV, JSON, Parquet). You will gain hands-on experience setting up isolated virtual environments, installing dependencies with pip, requesting data from APIs using httpx, and using standard debugging tools directly within VS Code and the CLI.Data Engineering with Python Foundations Bootcamp Only
-
Python Development & Testing
Transition from basic scripting to structured development workflows. You will learn to safely manage code changes using Git and GitHub (branching, commits, diffs, and merge flows) through VS Code. You will implement basic testing with pytest, track execution using logging (Loglyte), and structure state using Python classes. Through source-to-output reconciliation and data validation checks, you will learn to build repeatable scripts and answer the core engineering question: How do I know the data I made is the data I was supposed to make?Data Engineering with Python Foundations Bootcamp Only
-
Relational Databases & SQL Foundations
Build a solid foundation in database design, data modeling, and relational querying using PostgreSQL and Docker. You will write SQL queries to filter, aggregate, group, and join data across tables. You will learn to construct relational schemas using CREATE TABLE and Data Manipulation Language (DML), enforce primary/foreign keys and referential integrity, and programmatically integrate Python with PostgreSQL to validate and load structured datasets.Data Engineering with Python Foundations Bootcamp Only
-
The Data Engineering Lifecycle
Learn how raw data becomes a reliable asset across its entire lifecycle: ingestion, storage, transformation, orchestration, and delivery. You will gain hands-on experience making architectural design choices, evaluating compute and storage trade-offs, and embedding security, data quality validation, and observability into every step of the pipeline. -
Ingesting Data & Data Contracts
Pull structured and unstructured data from REST APIs, relational databases, flat files, and cloud object storage. You will adopt production engineering standards using modern project setup tooling (uv), dataclasses, and boundary validation with Pydantic. You will build resilient ingestion pipelines with pagination, authentication, retries, and idempotency while using explicit contracts to isolate malformed data into quarantine workflows. -
Data Storage & Lakehouse Architecture
Navigate storage selections based on analytical workload, access patterns, and scale. You will work across relational databases (PostgreSQL), embedded analytical engines (DuckDB), cloud object storage (AWS S3), and modern Data Lakehouse architectures using columnar formats (Parquet) and catalogs. You will evaluate partitioning strategies and learn when to leverage distributed processing engines like Apache Spark and Databricks. -
Data Transformation & Analytics Engineering
Convert raw source data into clean, business-ready analytical models. You will write advanced SQL using CTEs, window functions, and query execution plan analysis to optimize query performance. You will master analytics engineering with dbt, building modular transformation models, tracking automated data lineage, generating documentation, and enforcing data quality contracts with automated tests. -
Data Delivery & Event-Driven Systems
Deliver reliable datasets to downstream analytics platforms, machine learning models, and production endpoints. In addition to scheduled batch delivery, you will construct event-driven architectures using cloud object storage event listeners, message queues, topics, and consumer services. You will address streaming patterns, handle out-of-order event delivery, and manage low-latency data routing. -
Orchestration & Pipeline Resiliency
Transform isolated scripts into production-ready, automated workflows using Apache Airflow. You will design Directed Acyclic Graphs (DAGs), configure dependencies, schedule backfills, and implement failure handling. You will master observability, task state management, SLAs, and defensive checkpointing to ensure batch and event pipelines execute reliably without duplicate side effects. -
AI Engineering & Pipeline Integration
Treat Generative AI models as powerful but non-deterministic pipeline dependencies. Rather than basic prompt engineering, you will integrate local and hosted Large Language Models (LLMs) into data pipelines for structured extraction, document classification, and entity resolution. You will enforce strict output validation with Pydantic, optimize API costs via response/prompt caching and batching, track token telemetry, and design deterministic fallback mechanisms. -
Tools and Techniques
Work with the modern data engineering toolchain and software development practices. You will gain experience with VS Code, Git/GitHub, Docker, uv, pytest, dbt, Apache Airflow, DuckDB, PostgreSQL, Apache Spark/Databricks, and AWS S3. You will operate in agile team settings, conduct code reviews, document architectures, and communicate technical design trade-offs effectively. -
Real-World Projects & Capstone
Apply your learning across team-based engineering labs and an individual capstone project. Working with real-world datasets across healthcare, automotive, and public data domains, you will design, deploy, and defend an end-to-end data platform that demonstrates data quality validation, automated testing, observability, and cost-aware architecture. -
Career Preparation & Post-Graduation Support
Prepare to launch your career through direct career coaching and industry networking. You will interact with practicing data engineers, participate in technical resume and portfolio workshops, practice system design and coding interviews, and present your work to prospective employers at Demo Day. Post-graduation, you maintain access to dedicated job search support, community sessions, and continuing education seats.
Data Engineering with Python Foundations Bootcamp
-
Schedule
Tuesdays & Thursdays 6 - 9 pm CT -
Location
This class is live online (i.e. synchronous).
-
Tuition
$15,000
See below for detailed information on payment options.
Data Engineering Bootcamp
-
Schedule
Tuesdays & Thursdays 6 - 9 pm CT -
Location
This class is live online (i.e. synchronous).
-
Tuition
$12,500
See below for detailed information on payment options.