Skip to content

bench: validate rounded Score and isolated zero-fit results - #93

Open
Levius-Fubuki wants to merge 1 commit into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/bench-score-wire-rounding
Open

Levius-Fubuki wants to merge 1 commit into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/bench-score-wire-rounding

Conversation

@Levius-Fubuki

Copy link
Copy Markdown
Collaborator

Purpose

Decider rounds a derived Score to two decimals and its probabilities to four. The benchmark currently rejects a valid score=0.88 with probabilities 0.1234/0.8766 because it compares against 0.8766 with a fixed 1e-4 tolerance. Bound both independent rounding errors, use an accurate weighted sum, and require the score to remain inside the exact level range.

Also recognize the released isolated-Score zero-fit outcome only when its zero probabilities and complete fit/calibration metadata are consistent. Other invalid distributions still fail, and wire probabilities are preserved without renormalizing them. The net diff is one production Python file.

Test Plan

Baseline: 7f39ac40902c374803992407bb26eeba29c8a588 (main). Head: de34df45dd42f290c030b094fb630d8e65a82052.

  • python -m unittest discover -s tests/benchmarks -p 'test_*.py' -v
  • python benchmarks/bench.py validate benchmarks/smoke.jsonl
  • External regression harness: pinned legitimate Score and zero-fit responses; 1,152 deterministic independently rounded distributions across 2–10 levels; reversed label order; exact range; inconsistent/nonfinite values; incomplete/contradictory zero-fit metadata; huge integer metadata through actual request_one transport handling.
  • Full diff and independent review; git diff --check.

Test Result

8 existing tests and 4 smoke requests pass. All 10 external regressions pass, including the 1,152-distribution corpus. Before the change, the focused regression suite reproduced four validation errors and an out-of-range acceptance. Independent review found huge integer fit metadata could raise an uncaught overflow; equality guards now reject it before float conversion, and six end-to-end malformed-response cases are recorded as invalid_response.

Demo / evidence

CPU wire validation only; no GPU inference, numerical model parity, latency or throughput claim. Python 3.12.13, httpx 0.28.1, macOS ARM64. Reference: decider-ai 1.8.1 source, Decider-2B v11 revision. Scalar rounding bound is 0.005 + 0.00005 * n * (n-1) / 2 + 1e-9; probability-sum allowance from #75 is unchanged.

Regression scripts, fixtures, logs and reviewer reproductions remain local artifacts following the author's core-code-only diff preference. Existing repository tests are reproducible with the commands above; the additional local harness is not part of repository CI. Demo/video: N/A for a response validator.

Self-review

  • Reviewed the full diff and addressed findings.
  • Focused benchmark validation change; no serving/model architecture changes.
  • Ran relevant checks and reported their results and limitations.
  • Description and claims match the implementation and available evidence.

@Levius-Fubuki
Levius-Fubuki marked this pull request as ready for review October 5, 2026 16:19
Copilot AI balanced review requested due to automatic review settings October 5, 2026 16:19

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants