Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
42 changes: 42 additions & 0 deletions .github/workflows/portfolio-parse.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,42 @@
name: Parse portfolio dbt project

on:
pull_request:
branches: [main]
push:
branches: [main]
workflow_dispatch:

permissions:
contents: read

jobs:
parse:
runs-on: ubuntu-latest
timeout-minutes: 10
defaults:
run:
working-directory: task-2
env:
DATABRICKS_HOST: offline.invalid
DATABRICKS_HTTP_PATH: /sql/1.0/warehouses/offline
DATABRICKS_TOKEN: offline-placeholder
DBT_SCHEMA: portfolio_parse
DBT_SEND_ANONYMOUS_USAGE_STATS: 'false'
steps:
- uses: actions/checkout@v4
- uses: astral-sh/setup-uv@v6
with:
version: '0.8.22'
python-version: '3.12'
enable-cache: true
- name: Install original locked environment
run: uv sync --frozen
- name: Prepare dummy profile outside the checkout
run: |
mkdir -p "$RUNNER_TEMP/portfolio-dbt-profile"
cp profiles.yml.example "$RUNNER_TEMP/portfolio-dbt-profile/profiles.yml"
- name: Install dbt package and parse without warehouse access
run: |
uv run dbt deps --profiles-dir "$RUNNER_TEMP/portfolio-dbt-profile"
uv run dbt parse --profiles-dir "$RUNNER_TEMP/portfolio-dbt-profile" --no-partial-parse
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,7 @@ logs/
**/partial_parse.msgpack

# Python
.venv/
__pycache__/
*.py[cod]
*$py.class
Expand Down
62 changes: 56 additions & 6 deletions AI_ASSIST.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,13 +4,63 @@ Record at least one point where you used an AI coding assistant (ChatGPT, Claude

## Interaction 1

- **Tool used:** (e.g. ChatGPT / Cursor / Claude)
- **Task / Problem:** (e.g. debugging dbt connection profile / writing PySpark join / configuring Job trigger)
- **Tool used:** ChatGPT
- **Task / Problem:** Debugging the Databricks HTTP path being incorrectly converted by Git Bash.
- **Prompt sent:**
> `___`
> `i got this error

Connection:
00:21:33 host: **\*\*\*\***\*\***\*\*\*\***\***\*\*\*\***\*\***\*\*\*\***
00:21:33 http_path: C:/Program Files/Git/sql/1.0**\*\***\*\*\*\***\*\***
00:21:33 catalog: hyf
00:21:33 schema: dev_mohammedalfakih
00:21:33 Registered adapter: databricks=1.12.3
00:21:34 Databricks adapter: SPOG host (host_type='workspace'): no
00:21:34 Databricks adapter: workspace_id (from ?o= in http_path): None
00:21:34 Databricks adapter: databricks-sql-connector version: 4.3.0 (supported)
00:21:34 Databricks adapter: databricks-sdk version: 0.117.0 (supported)
00:21:34 databricks-sql-connector adapter: ThriftBackend.attempt_request: Exception: %s
ERROR:databricks.sql.backend.thrift_backend:ThriftBackend.attempt_request: Exception:
00:21:34 Databricks adapter: Connection(session-id=Unknown) - Exception while trying to create connection: Error during request to server.
Error properties: attempt=1/30, bounded-retry-delay=None, elapsed-seconds=0.10899758338928223/900.0, error-message=, http-code=404, method=OpenSession, no-retry-reason=non-retryable error, original-exception=, query-id=None, session-id=None
00:21:34 Connection test: [ERROR]

00:21:34 1 check failed:
00:21:34 dbt was unable to connect to the specified database.
The database returned the following error:

> Database Error
> Database Error

Error during request to server.

Check your database credentials and try again. For more information, visit:
https://docs.getdbt.com/docs/configure-your-profile

---

i notice that http path is loaded in wrong way. i have if in my .env as DATABRICKS_HTTP_PATH=/sql/1.0/warehouses/**\*\*\*\***\*\*\***\*\*\*\***
but in the error i can see http_path: C:/Program Files/Git/sql/1.0**\*\***\*\*\*\***\*\***
can u explain why is this happing and how to solve it.
note: the(\*) is only to hide secrets so I have an real values in my .env`

- **Output provided by AI:**
> `___`
> `You found the cause correctly. Git Bash uses MSYS2, which automatically converts environment-variable values that look like Unix paths before passing them to native Windows programs. Therefore:
/sql/1.0/warehouses/...
was incorrectly converted to:
C:/Program Files/Git/sql/1.0/warehouses/...
That produced the HTTP 404. MSYS2 documents this environment-variable path conversion.
Keep the value in .env unchanged—it is correct.
Next step only
In the same Git Bash session, run:
export MSYS2_ENV_CONV_EXCL=DATABRICKS_HTTP_PATH
uv run dbt debug --profiles-dir .
This tells Git Bash not to convert that specific environment variable. The debug output should now show an HTTP path beginning with:
/sql/1.0/warehouses/
and ideally finish with:
All checks passed!
`
- **What I kept, changed, or rejected, and why:**
> `___`
> `I kept the suggested `MSYS2_ENV_CONV_EXCL=DATABRICKS_HTTP_PATH`fix because it prevented Git Bash from converting the Databricks HTTP path, and`dbt debug`then passed. I kept the correct`/sql/1.0/warehouses/...` value unchanged and did not include any real token or credential in the AI prompt.`

*(Ensure no personal passwords, Databricks tokens, or unapproved credentials are included in prompts or logged outputs.)*
_(Ensure no personal passwords, Databricks tokens, or unapproved credentials are included in prompts or logged outputs.)_
131 changes: 69 additions & 62 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,74 +1,81 @@
# HackYourFuture Data Track — Week 13 Assignment

**Databricks Lab:** PySpark exploration, dbt incremental models on Delta Lake, and Git-backed Job scheduling.

Full instructions live in the curriculum: [Week 13 Assignment](https://www.notion.so/hackyourfuture/Assignment-2af50f64ffc98112b371c42a3f469749).

## Where to start

| Folder / File | What to submit | Points (autograder) |
| --- | --- | --- |
| `task-1/` | PySpark notebook (`show()` on aggregated results, PySpark-vs-dbt note) | 25 |
| `task-2/` | Ported dbt project + `WRITEUP.md` (timings + incremental explanation + `DESCRIBE HISTORY`) | 30 |
| `task-3/` | Git-backed Databricks Job + screenshots + `SCHEDULING.md` (Jobs vs Airflow) | 15 |
| `AI_ASSIST.md` | Documented LLM usage (prompt, tool, kept/discarded rationale) | 15 |
| Required files | Presence of all required files across `task-1/`, `task-2/`, `task-3/`, and `AI_ASSIST.md` | 15 |
| Secrets hygiene | No committed secrets (`profiles.yml`, `.env`, tokens) | Blocker if violated |

**Passing score:** 60/100 on the autograder. Your teacher also reviews quality against the rubric (incremental config, `>` boundary, tool-choice writing, Job configuration).

## Repository layout

```text
data-assignment-week-13/
├── task-1/
│ └── pyspark_exploration.ipynb # or .py export from Databricks
├── task-2/
│ ├── dbt_project.yml # your ported Week 10 project
│ ├── models/ # your dbt models
│ ├── profiles.yml.example
│ └── WRITEUP.md # incremental build write-up + DESCRIBE HISTORY
├── task-3/ # Git-backed Job scheduling
│ ├── SCHEDULING.md # Jobs vs Airflow write-up + Job Run URL
│ └── screenshots/ # Job config, green run, paused trigger
├── task-4/ # optional bonuses only (create if needed)
├── .env.example
├── AI_ASSIST.md # LLM interaction log
└── README.md
```
# NYC Taxi Databricks Lakehouse

## Setup
A Databricks lab exploring yellow-taxi data with PySpark, building a daily borough mart with dbt and Delta Lake, and running dbt from a Git-backed Databricks Job.

```bash
cp .env.example .env # fill in Databricks connection values
cd task-2
cp profiles.yml.example profiles.yml
export $(grep -v '^#' ../.env | xargs) # or source manually
dbt debug
**Questions:** Which pickup borough has the most records? How does the average total charge vary by payment type? How can a daily analytics table be updated with an incremental `MERGE`?

My completed HackYourFuture Data Track Week 13 work, with the original notebook, transformation logic and job evidence preserved. [Project provenance](docs/provenance.md) · [Original assignment guide](docs/hyf-assignment.md).

## Data flow

```mermaid
flowchart LR
R[Provided yellow-taxi raw trips] --> P[PySpark exploration]
Z[Provided taxi zone lookup] --> P
R --> T[dbt stg_trips view]
Z --> S[dbt stg_zones view]
T --> F[Delta daily borough mart]
S --> F
G[Git-backed Databricks Job] -. dbt deps + dbt build .-> F
F --> Q[Uniqueness, null and tip-ratio checks]
```

Use Python 3.11 or 3.12 for dbt if `dbt debug` crashes on import (`uvx --python 3.11 --from dbt-databricks dbt debug`).
**Stack:** Databricks · PySpark · SQL · dbt · Delta Lake · Databricks Jobs

## Explore the completed work

| Component | What to inspect |
|---|---|
| [PySpark notebook](task-1/pyspark_exploration.ipynb) | Zone join, borough counts, payment-type averages and saved aggregate outputs |
| [dbt project](task-2/) | Two staging views, the incremental daily mart, safe-division macro and data checks |
| [Incremental design](task-2/WRITEUP.md) | Merge key, date boundary, recorded timings and Delta history |
| [Job scheduling](task-3/SCHEDULING.md) | Git source, successful historical run, paused trigger and Jobs/Airflow comparison |

The notebook and dbt source definitions read `hyf.nyc_yellow.raw_trips` and `hyf.nyc_yellow.raw_zones`. Those normalized raw tables are provided by the course workspace; this repository does not ingest public TLC files or provision those tables. Their actual current coverage must be checked in the workspace.

### Saved notebook results

| Question | Recorded result |
|---|---|
| Highest pickup-borough count after the zone join | Manhattan: 112,028,489 records |
| Average `total_amount`, payment type 1 | $30.00 |
| Average `total_amount`, payment type 2 | $23.75 |

These are outputs saved in the original coursework notebook, not a new run or current city-wide statistics. Notebook queries use the raw population; the dbt mart applies staging filters, so their totals are not directly comparable.

## Check your score locally
### Historical job evidence

```bash
bash .hyf/test.sh
cat .hyf/score.json
![Successful Databricks coursework job run](task-3/screenshots/job_run_success.png)

The screenshot records a successful manual run on **30 July 2026**, including one incremental model, three passing data tests and one warning-level tip-ratio test. It is historical evidence, not a live deployment status. [Configuration, Git source and paused schedule screenshots](task-3/SCHEDULING.md).

## Model grain and incremental behavior

The model retains the assignment name **`fct_trips`**, but its grain is **one pickup borough per pickup date**. Its Delta merge key is `(pickup_borough, pickup_date)`; it is not a trip-level table. Keeping the name preserves the existing job command and coursework evidence.

It produces `trip_count`, `total_fare`, `avg_tip_pct` and `avg_trip_distance`. Staging excludes missing pickup IDs and negative/NULL fares; an inner zone join excludes unmatched pickup IDs. [Metric definitions and quality limits](docs/metrics.md).

The incremental filter selects dates strictly newer than the maximum date already in the target. This demonstrates an append-only daily load: late records, changes to prior days and additional records on the latest loaded day are skipped. A full refresh reprocesses them. [Current behavior and a future lookback design](task-2/WRITEUP.md).

## Setup and review

Requirements for a warehouse run: a Databricks workspace, SQL warehouse, access to the two normalized source tables, and permission to create views/tables in your own target schema. Python 3.12 and [uv](https://docs.astral.sh/uv/) install the existing locked local dbt environment.

From the repository root:

```sh
cd task-2
uv sync --frozen --python 3.12
```

## Scoring ladder (autograder)
Follow the [setup guide](docs/setup.md) to create an ignored local profile, set connection variables and build the upstream staging views plus the mart. An offline `dbt parse` review is also available without a Databricks account. Parsing validates project structure; it does not execute SQL or reproduce cloud results.

## Validation and scope

| Score | What the grader checks |
| --- | --- |
| 15 | Required files present (`task-1` notebook, `task-2/dbt_project.yml`, `WRITEUP.md`, `SCHEDULING.md`, `AI_ASSIST.md`) |
| 25 | Task 1 notebook mentions `show` and borough/payment_type work |
| 30 | Task 2 has incremental config (`materialized='incremental'`, `merge`, `unique_key`) and a filled `WRITEUP.md` |
| 15 | Task 3 has screenshots in `task-3/screenshots/`, Job Run URL, and a filled `SCHEDULING.md` |
| 15 | `AI_ASSIST.md` contains documented prompt and rationale |
| Pass | Secrets hygiene (no committed `.env` / `profiles.yml` / `dapi` tokens) |
Existing HYF checks are static coursework checks. The portfolio workflow additionally installs the lockfile and parses the dbt project with dummy connection settings. Neither CI path runs Databricks compute.

Governance and streaming bonuses are teacher-reviewed only; they do not affect the autograder score.
The original implementation is a learning project: source loading, late-data correction, production alerting and credential management are outside its scope. The job trigger was paused in the saved screenshot. The recorded full/incremental timings are a single historical pair, not a performance benchmark. [Validation scope](docs/validation.md).

## Instructor / track maintainer
## Credits

This repo is the Week 13 student scaffold. Teacher rubric: `week_13__assignment_rubric.md` in the [datatrack](https://github.com/HackYourFuture/datatrack) curriculum repo (not shared with students).
Author: [Mohammed Alfakih](https://github.com/mohammedalfakih-dev). Built through the [HackYourFuture](https://www.hackyourfuture.net/) Data Track using its [Week 13 assignment scaffold](https://github.com/HackYourAssignment/c55-data-week-13). Taxi data is supplied through the course's Databricks tables; public NYC taxi datasets are published by the [NYC Taxi & Limousine Commission](https://www.nyc.gov/site/tlc/about/tlc-trip-record-data.page).
74 changes: 74 additions & 0 deletions docs/hyf-assignment.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
# HackYourFuture Data Track — Week 13 Assignment

**Databricks Lab:** PySpark exploration, dbt incremental models on Delta Lake, and Git-backed Job scheduling.

Full instructions live in the curriculum: [Week 13 Assignment](https://www.notion.so/hackyourfuture/Assignment-2af50f64ffc98112b371c42a3f469749).

## Where to start

| Folder / File | What to submit | Points (autograder) |
| --- | --- | --- |
| `task-1/` | PySpark notebook (`show()` on aggregated results, PySpark-vs-dbt note) | 25 |
| `task-2/` | Ported dbt project + `WRITEUP.md` (timings + incremental explanation + `DESCRIBE HISTORY`) | 30 |
| `task-3/` | Git-backed Databricks Job + screenshots + `SCHEDULING.md` (Jobs vs Airflow) | 15 |
| `AI_ASSIST.md` | Documented LLM usage (prompt, tool, kept/discarded rationale) | 15 |
| Required files | Presence of all required files across `task-1/`, `task-2/`, `task-3/`, and `AI_ASSIST.md` | 15 |
| Secrets hygiene | No committed secrets (`profiles.yml`, `.env`, tokens) | Blocker if violated |

**Passing score:** 60/100 on the autograder. Your teacher also reviews quality against the rubric (incremental config, `>` boundary, tool-choice writing, Job configuration).

## Repository layout

```text
data-assignment-week-13/
├── task-1/
│ └── pyspark_exploration.ipynb # or .py export from Databricks
├── task-2/
│ ├── dbt_project.yml # your ported Week 10 project
│ ├── models/ # your dbt models
│ ├── profiles.yml.example
│ └── WRITEUP.md # incremental build write-up + DESCRIBE HISTORY
├── task-3/ # Git-backed Job scheduling
│ ├── SCHEDULING.md # Jobs vs Airflow write-up + Job Run URL
│ └── screenshots/ # Job config, green run, paused trigger
├── task-4/ # optional bonuses only (create if needed)
├── .env.example
├── AI_ASSIST.md # LLM interaction log
└── README.md
```

## Setup

```bash
cp .env.example .env # fill in Databricks connection values
cd task-2
cp profiles.yml.example profiles.yml
export $(grep -v '^#' ../.env | xargs) # or source manually
dbt debug
```

Use Python 3.11 or 3.12 for dbt if `dbt debug` crashes on import (`uvx --python 3.11 --from dbt-databricks dbt debug`).

## Check your score locally

```bash
bash .hyf/test.sh
cat .hyf/score.json
```

## Scoring ladder (autograder)

| Score | What the grader checks |
| --- | --- |
| 15 | Required files present (`task-1` notebook, `task-2/dbt_project.yml`, `WRITEUP.md`, `SCHEDULING.md`, `AI_ASSIST.md`) |
| 25 | Task 1 notebook mentions `show` and borough/payment_type work |
| 30 | Task 2 has incremental config (`materialized='incremental'`, `merge`, `unique_key`) and a filled `WRITEUP.md` |
| 15 | Task 3 has screenshots in `task-3/screenshots/`, Job Run URL, and a filled `SCHEDULING.md` |
| 15 | `AI_ASSIST.md` contains documented prompt and rationale |
| Pass | Secrets hygiene (no committed `.env` / `profiles.yml` / `dapi` tokens) |

Governance and streaming bonuses are teacher-reviewed only; they do not affect the autograder score.

## Instructor / track maintainer

This repo is the Week 13 student scaffold. Teacher rubric: `week_13__assignment_rubric.md` in the [datatrack](https://github.com/HackYourFuture/datatrack) curriculum repo (not shared with students).
20 changes: 20 additions & 0 deletions docs/metrics.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
# Model grain, metrics and quality

`fct_trips` contains one row per pickup borough and pickup date. Its composite merge key is `(pickup_borough, pickup_date)`. Despite the inherited model name, it is a daily borough aggregate, not a record of individual trips.

| Metric | Meaning |
|---|---|
| `trip_count` | Count of source records that pass staging and match a pickup zone |
| `total_fare` | Sum of meter fares in USD; not total payments or profit |
| `avg_tip_pct` | Mean of the implemented per-record tip/fare ratios; ratio rather than percent |
| `avg_trip_distance` | Mean recorded distance in miles |

Staging drops missing pickup IDs and negative/NULL fares. The zone join drops unmatched pickup IDs; a matching lookup row labeled `Unknown` still remains. Missing pickup timestamps are reported by a not-null test rather than filtered explicitly. There is no additional distance, tip or physical-trip deduplication policy in this lab.

The inherited `safe_divide` macro casts operands to `numeric` and uses `NULLIF` for zero denominators. On Databricks, bare `NUMERIC` defaults to `DECIMAL(10,0)`, so fractional dollar values can lose precision before division. The current SQL is preserved; a future correction should set an explicit decimal scale and validate fractional-cent/zero-denominator cases. Recorded tip ratios should not be treated as a newly verified exact-money calculation. [Databricks decimal type](https://docs.databricks.com/aws/en/sql/language-manual/data-types/decimal-type).

TLC's public yellow-taxi dictionary states that recorded tips exclude cash tips. The course tables are provided separately, so the public field definitions are context rather than independently verified provenance of every loaded row. [TLC yellow-taxi field dictionary](https://www.nyc.gov/assets/tlc/downloads/pdf/data_dictionary_trip_records_yellow.pdf).

The warning test returns groups whose mean tip ratio exceeds 1. The saved Job run returned 158 such groups with warning severity. A warning does not fail that run and does not prove the underlying values are correct or incorrect; it identifies groups to investigate.

The PySpark notebook answers different raw-data questions. Its borough query joins raw trips to zones without the dbt fare filter; its payment query averages raw `total_amount`. These outputs must not be equated with the mart's fare sum or trip counts. When combining daily rows, sum counts/fares; do not average daily averages without the relevant contributing counts.
Loading
Loading