bench: accept wire-rounded probability sums per label count - #75
Merged
Merged
Conversation
The validator's sum check used a flat abs_tol=1e-4, which rejected four-value distributions that the worker's four-decimal rounding makes sum to 0.9999 (PR #40's first GPU run, comment 5894999739: two MMLU responses rejected at the floating-point boundary, stopping the run). The allowance now derives from the label count: each rounded value carries up to half a decimal step, and the sum is computed with math.fsum. A regression test replays both observed vectors and checks that a genuinely off distribution is still rejected.
4 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Fix the live validator bug in
benchmarks/bench.pyfound in #40's first GPU run(comment):
the probability sum check used a flat
math.isclose(..., abs_tol=1e-4), but workers roundprobabilities to four decimal places on the wire, so a rounded distribution can legitimately
miss 1 by more than 1e-4. Two MMLU responses summing to
0.9998999999999999were rejected atthe floating-point boundary, which stopped the first measured GPU run after 498/500 requests.
The tolerance is now derived from the wire format instead of being a flat constant: each of the
nrounded values carries up to half a decimal step (5e-5), so the acceptance bound isn * 5e-5(plus epsilon for float representation), and the sum is computed withmath.fsumto remove addition-order noise. A four-label distribution summing to 0.9999 is accepted; a
distribution off by a centi-level (e.g. 0.99) is still rejected, so this widens acceptance
only by what the rounding itself can produce.
The score-vs-expected-value check keeps its
abs_tol=1e-4: no failure of that check has beenobserved, and its deviation sources (worker-side rounding of a derived scalar) are different.
Test Plan
System1-Omni Version / Commit:
mainat69c3485(after #67)(
[0.1099, 0.1211, 0.5589, 0.2100],[0.1130, 0.5135, 0.3151, 0.0583]) throughread_answers, and asserts a four-value distribution summing to 0.99 is still rejected.ValueError: probabilities must sum to oneand passes with the fix.python -m unittest discover -s tests/benchmarks -p 'test_*.py'— 8 tests OK(Python 3.13, httpx 0.28.1; CI's 3.11 is equivalent here).
python benchmarks/bench.py validate benchmarks/smoke.jsonl— OK, unchanged SHA256.Test Result
All weight-free checks green locally on macOS (arm64). No benchmark numbers are claimed by
this PR; it changes validation only, and it un-blocks judged parity runs for the Laya native
engine chain by no longer rejecting legitimate rounded responses.
Self-review