perf: speed up filtering and graph data preparation for large datasets (#364) - #365
Open
timdegroot1996 wants to merge 2 commits into
Open
timdegroot1996 wants to merge 2 commits into
timdegroot1996 wants to merge 2 commits into
Conversation
#364) Applying a filter on a large dataset spent most of its time in our own JavaScript rather than in Chart.js: - filter_data looked up every suite/test/keyword row in an array of run starts (rows x runs); it now uses a Set. - The run_start transformations (milliseconds, timezone conversion, timezone display) copied every row on every filter apply. They only depend on three settings, so the copies are cached per combination and shared with the filter option availability, which had its own cache. - The timeline views of Most Failed, Most Flaky and Messages scanned all rows for every run and label. Rows are now grouped once by label and run_start. With a message config every rule's regex is built once. - The suite section filters were read from the DOM for every row; get_suite_data_exclusion reads them once and returns the check. - sort_wall_clock computes its keys once, the test select no longer does a quadratic duplicate check, and the base64 payload is decoded with a plain loop instead of Uint8Array.from with a callback. On a 504 run / 208k test dataset applying the default filter goes from 1.18 s to 0.37 s, all runs from 8.1 s to 3.1 s, and the initial load from 3.2 s to 2.3 s. Chart data, filtered data and select options are identical to before across timestamp settings, amounts, bar/timeline views, suite paths and a message config. The message config timeline also no longer relies on a block-scoped function and an undeclared variable, which threw in strict mode. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ion (#364) - Overview: every run card created its own Chart.js donut, twice, on every update, and the donuts of removed cards were never destroyed, so Chart.js kept all of them alive. A donut is now created when its card comes near the viewport and destroyed before its card is removed. Updating the overview with 504 runs goes from 4.0 s to 0.07 s. - Tables: DataTables detected the type of every column by checking every cell on every update. The columns now get the type it detected, except the version which depends on the project. 3.9 s to 2.4 s. - Data loading: the embedded payloads are inflated in parallel with the native DecompressionStream instead of pako, which is removed. main() awaits load_data() before anything reads the data, so the transform cache in the filter pipeline reads the arrays when it is called. - build_tooltip_meta parses every run_start once instead of per row. The initial load of the 504 run dashboard goes from 3.8 s to 2.3 s. Co-Authored-By: HuntTheSun <53446567+HuntTheSun@users.noreply.github.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #364
Problem
On large dashboards, applying a filter spent most of its time in our own JavaScript rather than in Chart.js. On a 504-run / 208k-test dataset, filtering took 40–66% of the time to apply a filter. #364 has the full profiling.
Changes
filter_data: looks up run starts in aSetinstead of scanning an array for every suite, test and keyword row (rows × runs).group_timeline_values) instead of scanning all rows for every run and label. With a message config, each rule's regex is built once.get_suite_data_exclusionreads the selected suite filters from the DOM once and returns the per-row check.exclude_from_suite_dataused to read them for every row.sort_wall_clockcomputes its sort keys once.Uint8Array.fromwith a callback.Results
504 runs / 208k tests, headless Chrome, median of 3:
Chart.js is the largest remaining cost. The follow-up items are listed in #364.
Tests
mainand this branch and compared chart data, filtered data and select options:group_timeline_valuesget_suite_data_exclusion(the four selection cases, plus reading the DOM only once)Second commit: Overview, Tables, data loading
versionis left to detection. Sort and search results are identical tomain. Update: 3.9 s → 2.4 s.DecompressionStream(based on @HuntTheSun's branch), and pako is removed.main()awaitsload_data()before anything reads the data.build_tooltip_metaparses each run start once (180 → 51 ms).DecompressionStreamneeds Chrome 80+, Firefox 113+ or Safari 16.4+.🤖 Generated with Claude Code