Skip to content

feat: Add ability to generate more realistic learner journey data - #276

Open
bmtcril wants to merge 2 commits into
mainfrom
perf/engagement-harness
Open

bmtcril wants to merge 2 commits into
mainfrom
perf/engagement-harness

Conversation

@bmtcril

@bmtcril bmtcril commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Adds a journeys command that generates learner data with known expected results, so Aspects report numbers can be checked against a known answer instead of only being load-tested.

The existing backends pick each event independently. This new mode simulates each enrolled learner moving through a properly nested course (sections > subsections > units > problems and videos) in time-ordered sessions. The simulator knows exactly what each learner did, so it also writes the expected engagement results.

What it generates (gzipped CSV, same layouts as the csv backend):

  • Event sink rows: courses, blocks (several publishes, with deleted units only in earlier publishes), external ids and user profiles.
  • xAPI events matching event-routing-backends shapes:
    • navigation events whose object is the unit being left, as in the Learning MFE;
    • problem attempted / evaluated;
    • video initialized / played / paused / seeked / completed / terminated, with completed fired once at 95%;
    • optional mailto: actors.
  • expected_engagement: per learner and section / subsection, pages / problems / videos done out of total, with the Aspects status label.
  • expected_video_seconds: watched seconds per learner and video. "Observable" is what the events can show (each play paired with the next video event). "Truth" is what the learner actually watched, including tab closes.
  • manifest.json: counts, plus the most active learner per course for the learner dashboard.

Edge cases (all on in journeys_oracle.yaml): mailto: actors, deleted units, duplicate events, pause and resume at the same instant, recent activity inside a refreshable-view lookback window, and still-running courses.

Example configs: journeys_oracle.yaml (~9K events, for correctness), journeys_m.yaml (~15M) and journeys_l.yaml (~75M) for benchmarks. The same config, seed and --now produce identical files.

Testing

  • 18 tests in test_journeys.py. The main check rebuilds every expected page / problem row and every observable watched-seconds value from the generated xAPI events and block structure alone, so the simulator's bookkeeping can't drift from the events it emits. The others cover reproducibility, ordering, nesting, edge-case presence, video pairing edge cases, the CLI and config errors.
  • Used to benchmark and check correctness of aspects-dbt's engagement models. The oracle and M datasets load into a ClickHouse 25.8 / Aspects stack and parse through the full MV chain.

Usage

xapi-db-load journeys --config_file example_configs/journeys_oracle.yaml --now "2026-10-08 12:00:00"

Loading instructions are in the README.

🤖 Generated with Claude Code

@codecov

codecov Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.41935% with 16 lines in your changes missing coverage. Please review.
✅ Project coverage is 88.37%. Comparing base (9f7cac1) to head (457b3c4).

Files with missing lines Patch % Lines
xapi_db_load/journeys/simulate.py 94.84% 5 Missing and 8 partials ⚠️
xapi_db_load/journeys/structure.py 97.84% 1 Missing and 1 partial ⚠️
xapi_db_load/journeys/generate.py 99.25% 0 Missing and 1 partial ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #276      +/-   ##
==========================================
+ Coverage   85.54%   88.37%   +2.82%     
==========================================
  Files          29       34       +5     
  Lines        1986     2606     +620     
  Branches      175      270      +95     
==========================================
+ Hits         1699     2303     +604     
- Misses        260      266       +6     
- Partials       27       37      +10     
Flag Coverage Δ
unittests 88.37% <97.41%> (+2.82%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@bmtcril
bmtcril marked this pull request as draft October 9, 2026 17:36
- Stop counting a video's actual watched seconds past "now" for sessions still playing.
- Report journeys config problems as CLI usage errors.
- Check course_length_days fits in window_days.
- Add tests for the CLI, video pairing edge cases and unsorted output.
- Document the journeys command in the README; fix dataset size comments.
- Lint fixes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant