Repository navigation
fix(apa-57): never report QUALIFIED when the LLM matrix measured nothing - #123
Conversation
The anti-fake-green gap that let an unmeasured run report success.
Observed while attempting the APA-56 causal re-measurement: every live case
self-skipped on provider 429, so zero model measurements were recorded, and
qualify.sh still printed
QUALIFIED: full suite green against a genuinely fresh, ephemeral database
with exit 0. The existing anti-fake-green check only proves PG-backed tests
RAN; it cannot see a live matrix, so green-by-skipping was indistinguishable
from a real qualification. A verdict is what a ledger records, so a wrong
headline is worse than no gate at all.
The Go harness was honest throughout — it Skipf'd with "the real-LLM matrix
cannot be measured right now. This is a credential/provider condition, not a
result" — which is the only reason this was caught. The harness is left
exactly as it is; the reporting was what lied.
A run now declares its matrix with QUAL_LLM_SCENARIOS, and then:
requested > 0 AND measured == 0 -> INFRA_BLOCKED, non-zero exit
0 < measured < expected -> INCOMPLETE, non-zero exit
measured == expected -> eligible for a verdict
A scenario counts as measured only if the harness emitted an APA49-EVIDENCE
record for it. A 429, a skip, a discard, or a fatal is never a measurement —
those are precisely the outcomes that used to produce a green verdict with
nothing behind it. An incomplete run names the scenarios it did not measure,
so measured and unmeasured are always reported separately.
Declaring no scenarios leaves the PG-only path unchanged and independently
valid; its verdict line now says so explicitly rather than implying a model
measurement that never happened.
The verbose go test output is now tee'd instead of discarded. A record that
is never captured cannot be counted, so accounting requires capture.
No change to model, prompt, decoder, validator, or any production behaviour.
MaxTokens is untouched, the prompt is untouched, and APA-55 stays uncommitted.
Verified: bash -n clean. PG-only run green, exit 0, 20 PG-backed tests. The
0/N branch reproduced against the real provider 429 condition: exit 1,
QUALIFIED suppressed, all 6 scenarios named. Partial and complete branches
verified with a throwaway stub standing in for `go test` (3/6 -> INCOMPLETE
exit 1; 6/6 -> QUALIFIED exit 0); the stub was not committed, because the repo
has no shell-test convention and adding one would widen this change.
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configuration
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Aparnap2
left a comment
There was a problem hiding this comment.
Reviewed PR #123 against the anti-fake-green failure mode and the qualification contract. PASS / merge-ready.
Verified from the diff: (1) the gate is opt-in via QUAL_LLM_SCENARIOS, so PG-only qualification remains independently valid; (2) measurement is correctly tied to emitted APA49-EVIDENCE records, so 429/skip/discard/fatal cannot count; (3) 0/expected is INFRA_BLOCKED and partial is INCOMPLETE, both non-zero; (4) the full matrix is required before an LLM qualification can return 0; (5) stdout is captured with tee while pipefail preserves go-test status; (6) the declared scenario names are reported when unmeasured; and (7) no model/prompt/decoder/validator/MaxTokens behavior is changed.
The throwaway-stub verification is a test-coverage limitation, not a merge blocker. A durable shell regression test can be a separate follow-up if needed. CI being UNSTABLE is not caused by this diff based on the current commit status; the only reported check here is CodeRabbit success, and the PR notes that qualify.sh is not invoked by CI.
The hole this closes
Discovered while attempting the APA-56 causal re-measurement:
infra/postgres/qualify.shprintedQUALIFIEDwith exit 0 on a run where no model behaviour was measured at all.The existing anti-fake-green check only proves PG-backed tests ran (
grep -c '^--- PASS'on./internal/repository/postgres/). It has no equivalent for a live matrix, so green-by-skipping was indistinguishable from a real qualification.The Go harness was honest throughout — it
t.Skipf'd with "the real-LLM matrix cannot be measured right now. This is a credential/provider condition, not a result". That honesty is the only reason this was caught. The harness is untouched; the reporting was what lied.The gate
A run declares its matrix with
QUAL_LLM_SCENARIOS, and then:INFRA_BLOCKEDINCOMPLETEA scenario counts as measured only if the harness emitted an
APA49-EVIDENCErecord for it. A 429, skip, discard, or fatal is never a measurement — those are exactly the outcomes that used to produce a green verdict with nothing behind it.An incomplete run names the scenarios it did not measure, so measured and unmeasured are always reported separately:
Verified
bash -nQUALIFIED, 20 PG-backed tests ranQUALIFIEDsuppressed, all 6 namedgo testINCOMPLETE, 3 namedgo testQUALIFIEDThe 0/6 branch was reproduced against the actual failure condition, not simulated. Partial/complete used a throwaway stub standing in for
go test— not committed, because the repo has no shell-test convention and adding one would widen this change. Those two branches will be exercised by the next real run once provider capacity allows.Scope discipline
MaxTokens: 4096untouched; the prompt is not shrunk to fit a quota.claimops-postgres(shared) never contacted; teardown left noclaimops-pg-qualcontainer.One supporting change: the verbose
go testoutput is nowtee'd instead of discarded. A record that is never captured cannot be counted, so accounting requires capture.Closes APA-57. Blocks the APA-55 causal re-measurement until provider capacity is resolved.