Skip to content
View MasouData's full-sized avatar

Block or report MasouData

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
MasouData/README.md

Hi, I'm Masoud Aghayan 👋

Data Engineer & Business Data Analyst
Databricks • PySpark • SQL • Python • Azure • Data Analytics

Based in Utrecht, Netherlands 🇳🇱

I build reliable data pipelines, analytics solutions and applied machine-learning systems that turn complex data into useful business decisions.

My background combines data engineering, banking and risk analytics, industrial data at ASML, and applied data science. I completed an MSc in Computer Science – Data Science at Utrecht University in 2025 and hold the Databricks Certified Data Engineer Associate certification.

What I work with

Area Technologies & Methods
Data Engineering Python, PySpark, SQL, Apache Spark, Databricks, Delta Lake, Kafka, Azure Data Factory
Data Platforms Lakehouse architecture, Bronze/Silver/Gold, batch & streaming pipelines, data quality
DataOps / MLOps Git, GitHub Actions, Databricks Asset Bundles, MLflow, Unity Catalog
Analytics & ML scikit-learn, XGBoost, statistical modelling, experimentation, explainable ML
AI Applications LangGraph, RAG, Databricks Vector Search, Microsoft Foundry
BI & Reporting Power BI, Spotfire

Featured Projects

🔄 Real-Time Kafka → Databricks Lakehouse

End-to-end streaming pipeline using Kafka, PySpark, Databricks Lakeflow and Delta Lake, with Bronze–Silver–Gold processing, data-quality rules, orchestration and dev/prod deployment.

Tech: Kafka · PySpark · Databricks · Delta Lake · Lakeflow · Power BI
View project


🏦 Explainable KYC Risk Modelling

Customer-review prioritisation using 5.08M synthetic transactions. Compares a transparent scorecard, Logistic Regression and Explainable Boosting Machine using calibration, precision, recall and lift.

Tech: Python · PySpark · Databricks · Interpretable ML
View project


🤖 Explainable Credit Risk Copilot

Governed credit-risk decision-support prototype combining statistical models, deterministic policy rules, document retrieval and generative AI.

Tech: Databricks · MLflow · Vector Search · LangGraph · Microsoft Foundry
View project


🇳🇱 Housing Energy Behaviour in the Netherlands

Reproducible analysis of Dutch CBS open data across 437 municipalities, combining data engineering, PySpark and panel-data econometrics.

Tech: Python · PySpark · R · CBS StatLine · Econometrics
View project


🚀 Databricks CI/CD with Asset Bundles

Automated Databricks deployment workflow using GitHub Actions, Databricks Asset Bundles and separate development/production targets.

Tech: Databricks · GitHub Actions · YAML · CI/CD · DataOps
View project


📊 Mobile Marketing Attribution & Customer Lifetime Value

End-to-end business analytics project covering acquisition performance, attribution, customer economics and predictive 180-day LTV.

Tech: Python · Analytics · Machine Learning · Customer Economics
View project

Experience Highlights

ASML — Data Engineer / Data Scientist Intern
Worked with large operational datasets using Python, SQL, PySpark and Databricks, including predictive modelling and operational analytics.

Banking — Financial Data Analyst
Five years working with customer, transaction, KYC/CDD, credit-risk and portfolio data.

Omdena – TerraYield 2 — Execution Anchor & Data Engineer Collaborator
Working with heterogeneous satellite, weather, environmental and economic data in an international project.

Education & Certification

MSc Computer Science – Data Science, Utrecht University, 2025
Thesis: Data Valuation & Reduction in Dataset Federations

Databricks Certified Data Engineer Associate, 2026

Connect

LinkedInGitHub

Pinned Loading

  1. databricks-asset-bundles-ci-cd databricks-asset-bundles-ci-cd Public

    CI/CD for Databricks using Asset Bundles and GitHub Actions with dev/prod targets, automated deployment and job execution

    Jupyter Notebook

  2. explainable-credit-risk-copilot explainable-credit-risk-copilot Public

    Governed credit-risk decision-support prototype using Databricks, MLflow, Vector Search, LangGraph and Microsoft Foundry.

    Jupyter Notebook

  3. explainable-kyc-risk-modelling explainable-kyc-risk-modelling Public

    Explainable KYC periodic-review prioritisation on 5.08M synthetic transactions using PySpark, Logistic Regression and Explainable Boosting Machines.

    Python

  4. housing-energy-behaviour-nl housing-energy-behaviour-nl Public

    Housing Energy Behavior in Netherlands: Retrofit, Demand and Equity

    Python

  5. mobile-marketing-attribution-customer-lifetime-value mobile-marketing-attribution-customer-lifetime-value Public

    An end-to-end analysis of mobile acquisition performance, privacy-safe attribution, and predictive Customer Lifetime Value (LTV) modeling.

    Python

  6. real-time-kafka-databricks-pipeline real-time-kafka-databricks-pipeline Public

    End-to-end streaming lakehouse with Kafka, PySpark, Databricks Lakeflow, Delta Lake, Asset Bundles and Power BI.

    Python