Data Engineer & Business Data Analyst
Databricks • PySpark • SQL • Python • Azure • Data Analytics
Based in Utrecht, Netherlands 🇳🇱
I build reliable data pipelines, analytics solutions and applied machine-learning systems that turn complex data into useful business decisions.
My background combines data engineering, banking and risk analytics, industrial data at ASML, and applied data science. I completed an MSc in Computer Science – Data Science at Utrecht University in 2025 and hold the Databricks Certified Data Engineer Associate certification.
| Area | Technologies & Methods |
|---|---|
| Data Engineering | Python, PySpark, SQL, Apache Spark, Databricks, Delta Lake, Kafka, Azure Data Factory |
| Data Platforms | Lakehouse architecture, Bronze/Silver/Gold, batch & streaming pipelines, data quality |
| DataOps / MLOps | Git, GitHub Actions, Databricks Asset Bundles, MLflow, Unity Catalog |
| Analytics & ML | scikit-learn, XGBoost, statistical modelling, experimentation, explainable ML |
| AI Applications | LangGraph, RAG, Databricks Vector Search, Microsoft Foundry |
| BI & Reporting | Power BI, Spotfire |
End-to-end streaming pipeline using Kafka, PySpark, Databricks Lakeflow and Delta Lake, with Bronze–Silver–Gold processing, data-quality rules, orchestration and dev/prod deployment.
Tech: Kafka · PySpark · Databricks · Delta Lake · Lakeflow · Power BI
View project
Customer-review prioritisation using 5.08M synthetic transactions. Compares a transparent scorecard, Logistic Regression and Explainable Boosting Machine using calibration, precision, recall and lift.
Tech: Python · PySpark · Databricks · Interpretable ML
View project
Governed credit-risk decision-support prototype combining statistical models, deterministic policy rules, document retrieval and generative AI.
Tech: Databricks · MLflow · Vector Search · LangGraph · Microsoft Foundry
View project
Reproducible analysis of Dutch CBS open data across 437 municipalities, combining data engineering, PySpark and panel-data econometrics.
Tech: Python · PySpark · R · CBS StatLine · Econometrics
View project
Automated Databricks deployment workflow using GitHub Actions, Databricks Asset Bundles and separate development/production targets.
Tech: Databricks · GitHub Actions · YAML · CI/CD · DataOps
View project
End-to-end business analytics project covering acquisition performance, attribution, customer economics and predictive 180-day LTV.
Tech: Python · Analytics · Machine Learning · Customer Economics
View project
ASML — Data Engineer / Data Scientist Intern
Worked with large operational datasets using Python, SQL, PySpark and
Databricks, including predictive modelling and operational analytics.
Banking — Financial Data Analyst
Five years working with customer, transaction, KYC/CDD, credit-risk and
portfolio data.
Omdena – TerraYield 2 — Execution Anchor & Data Engineer Collaborator
Working with heterogeneous satellite, weather, environmental and economic
data in an international project.
MSc Computer Science – Data Science, Utrecht University, 2025
Thesis: Data Valuation & Reduction in Dataset Federations
Databricks Certified Data Engineer Associate, 2026
