Skip to content

bench: accept wire-rounded probability sums per label count - #75

Merged
hsliuustc0106 merged 1 commit into
mainfrom
fix/bench-probability-sum-tolerance
Oct 4, 2026
Merged

hsliuustc0106 merged 1 commit into
mainfrom
fix/bench-probability-sum-tolerance

Conversation

@hsliuustc0106

Copy link
Copy Markdown
Contributor

Purpose

Fix the live validator bug in benchmarks/bench.py found in #40's first GPU run
(comment):
the probability sum check used a flat math.isclose(..., abs_tol=1e-4), but workers round
probabilities to four decimal places on the wire, so a rounded distribution can legitimately
miss 1 by more than 1e-4. Two MMLU responses summing to 0.9998999999999999 were rejected at
the floating-point boundary, which stopped the first measured GPU run after 498/500 requests.

The tolerance is now derived from the wire format instead of being a flat constant: each of the
n rounded values carries up to half a decimal step (5e-5), so the acceptance bound is
n * 5e-5 (plus epsilon for float representation), and the sum is computed with math.fsum
to remove addition-order noise. A four-label distribution summing to 0.9999 is accepted; a
distribution off by a centi-level (e.g. 0.99) is still rejected, so this widens acceptance
only by what the rounding itself can produce.

The score-vs-expected-value check keeps its abs_tol=1e-4: no failure of that check has been
observed, and its deviation sources (worker-side rounding of a derived scalar) are different.

Test Plan

System1-Omni Version / Commit: main at 69c3485 (after #67)

  • New regression test replays the two exact vectors observed in the GPU run
    ([0.1099, 0.1211, 0.5589, 0.2100], [0.1130, 0.5135, 0.3151, 0.0583]) through
    read_answers, and asserts a four-value distribution summing to 0.99 is still rejected.
  • Verified the regression test fails on the pre-fix code with the original
    ValueError: probabilities must sum to one and passes with the fix.
  • python -m unittest discover -s tests/benchmarks -p 'test_*.py' — 8 tests OK
    (Python 3.13, httpx 0.28.1; CI's 3.11 is equivalent here).
  • python benchmarks/bench.py validate benchmarks/smoke.jsonl — OK, unchanged SHA256.

Test Result

All weight-free checks green locally on macOS (arm64). No benchmark numbers are claimed by
this PR; it changes validation only, and it un-blocks judged parity runs for the Laya native
engine chain by no longer rejecting legitimate rounded responses.

Self-review

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

The validator's sum check used a flat abs_tol=1e-4, which rejected
four-value distributions that the worker's four-decimal rounding makes
sum to 0.9999 (PR #40's first GPU run, comment 5894999739: two MMLU
responses rejected at the floating-point boundary, stopping the run).
The allowance now derives from the label count: each rounded value
carries up to half a decimal step, and the sum is computed with
math.fsum. A regression test replays both observed vectors and checks
that a genuinely off distribution is still rejected.
@hsliuustc0106
hsliuustc0106 merged commit 579fa8f into main Oct 4, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant