Turning coffee into data pipelines!
I'm an early-career Data Engineer based in Tampere, Finland, currently completing an MSc in Computer Science & Electrical Engineering at Tampere University.
I enjoy building data pipelines that are clean, reliable, and useful — from raw data ingestion to transformation, storage, and analysis.
Currently focused on:
- Batch and streaming pipelines with PySpark
- Data lake formats such as Delta Lake and Parquet
- Workflow orchestration with Airflow
- Databricks-based data engineering projects
- Writing cleaner, more testable data code
Languages: Python · SQL · Java
Data Engineering: PySpark · Spark Structured Streaming · ETL/ELT · Delta Lake · Parquet · Airflow · dbt
Platforms & Tools: Databricks · Azure Blob Storage · Docker · Git · PostgreSQL · MySQL
Benchmarked CSV, Parquet, and Delta Lake using PySpark and Databricks, comparing storage format behavior, query performance, and schema evolution.
Tech: PySpark · Delta Lake · Databricks · Azure Blob Storage
Built a PySpark pipeline combining structured streaming with machine learning to process and analyze retail sales data.
Tech: PySpark · Spark Structured Streaming · Python
Built a production-inspired pipeline that ingests weather data from the Open-Meteo API, orchestrates ETL workflows with Airflow, and models it into a star schema in PostgreSQL for analytics.
Tech: Python · Apache Airflow · PostgreSQL · dimensional modeling
Associate Data Engineer — DataCamp, 2026
MSc, Computer Science & Electrical Engineering
Tampere University, Finland
BSc, Information Systems & Technology
Makerere University
- Portfolio: hamza-sentongo.netlify.app
- LinkedIn: https://www.linkedin.com/in/hamza-sentongo-55483a243
