bench: validate rounded Score and isolated zero-fit results - #93
Open
Levius-Fubuki wants to merge 1 commit into
Open
Levius-Fubuki wants to merge 1 commit into
Levius-Fubuki wants to merge 1 commit into
Conversation
Levius-Fubuki
marked this pull request as ready for review
October 5, 2026 16:19
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Decider rounds a derived Score to two decimals and its probabilities to four. The benchmark currently rejects a valid
score=0.88with probabilities0.1234/0.8766because it compares against0.8766with a fixed1e-4tolerance. Bound both independent rounding errors, use an accurate weighted sum, and require the score to remain inside the exact level range.Also recognize the released isolated-Score zero-fit outcome only when its zero probabilities and complete fit/calibration metadata are consistent. Other invalid distributions still fail, and wire probabilities are preserved without renormalizing them. The net diff is one production Python file.
Test Plan
Baseline:
7f39ac40902c374803992407bb26eeba29c8a588(main). Head:de34df45dd42f290c030b094fb630d8e65a82052.python -m unittest discover -s tests/benchmarks -p 'test_*.py' -vpython benchmarks/bench.py validate benchmarks/smoke.jsonlrequest_onetransport handling.git diff --check.Test Result
8 existing tests and 4 smoke requests pass. All 10 external regressions pass, including the 1,152-distribution corpus. Before the change, the focused regression suite reproduced four validation errors and an out-of-range acceptance. Independent review found huge integer fit metadata could raise an uncaught overflow; equality guards now reject it before float conversion, and six end-to-end malformed-response cases are recorded as
invalid_response.Demo / evidence
CPU wire validation only; no GPU inference, numerical model parity, latency or throughput claim. Python 3.12.13, httpx 0.28.1, macOS ARM64. Reference: decider-ai 1.8.1 source, Decider-2B v11 revision. Scalar rounding bound is
0.005 + 0.00005 * n * (n-1) / 2 + 1e-9; probability-sum allowance from #75 is unchanged.Regression scripts, fixtures, logs and reviewer reproductions remain local artifacts following the author's core-code-only diff preference. Existing repository tests are reproducible with the commands above; the additional local harness is not part of repository CI. Demo/video: N/A for a response validator.
Self-review