Data Engineering
Build the pipelines that feed every decision.
Data engineers move, clean and organise data so analysts and ML teams can use it.
What the work looks like
Building ETL pipelines, designing warehouses, fixing data quality issues, optimising queries.
Key skills
- SQL (advanced)
- Python
- Data modelling
- Distributed processing
- Pipeline orchestration
Common tools
SparkAirflowdbtSnowflake / BigQueryKafka
Pros
- Fast-growing demand
- Less DSA-heavy interviews
- Foundation for ML and analytics
Cons
- Pipelines break at 3am
- Less visible than data science
- Messy real-world data
A 12-month starter roadmap
Months 1–3
- • SQL deeply
- • Python
- • Git
- • Linux basics
Months 4–6
- • Data modelling
- • Airflow
- • A cloud warehouse
- • Build one pipeline
Months 7–12
- • Spark
- • Streaming with Kafka
- • dbt
- • Cloud certification (optional)
Reality check
Students like: Making chaotic data reliable.
Students dislike: Broken upstream data you don't control.
Hardest part: Designing for scale and correctness.
Common myth: It's not 'just SQL' — it's distributed systems work.