Skip to content

fix(apa-57): never report QUALIFIED when the LLM matrix measured nothing - #123

Merged
Aparnap2 merged 1 commit into
mainfrom
fix/apa-57-llm-anti-fake-green-gate
Oct 6, 2026
Merged

Aparnap2 merged 1 commit into
mainfrom
fix/apa-57-llm-anti-fake-green-gate

Conversation

@Aparnap2

@Aparnap2 Aparnap2 commented Oct 6, 2026

Copy link
Copy Markdown
Owner

The hole this closes

Discovered while attempting the APA-56 causal re-measurement:

LLM matrix
   ↓
every live case self-skips (provider 429)
   ↓
0 model measurements recorded
   ↓
PG suite passes
   ↓
qualify.sh → "QUALIFIED" · exit 0

infra/postgres/qualify.sh printed QUALIFIED with exit 0 on a run where no model behaviour was measured at all.

The existing anti-fake-green check only proves PG-backed tests ran (grep -c '^--- PASS' on ./internal/repository/postgres/). It has no equivalent for a live matrix, so green-by-skipping was indistinguishable from a real qualification.

The Go harness was honest throughout — it t.Skipf'd with "the real-LLM matrix cannot be measured right now. This is a credential/provider condition, not a result". That honesty is the only reason this was caught. The harness is untouched; the reporting was what lied.

The gate

A run declares its matrix with QUAL_LLM_SCENARIOS, and then:

Condition Verdict Exit
requested > 0 AND measured == 0 INFRA_BLOCKED non-zero
0 < measured < expected INCOMPLETE non-zero
measured == expected eligible for a verdict 0

A scenario counts as measured only if the harness emitted an APA49-EVIDENCE record for it. A 429, skip, discard, or fatal is never a measurement — those are exactly the outcomes that used to produce a green verdict with nothing behind it.

An incomplete run names the scenarios it did not measure, so measured and unmeasured are always reported separately:

LLM cases measured: 3/6; skipped subtests: 0
LLM cases NOT measured: apa55_a5_cross_tenant apa55_a7_stale apa55_b1_valid_tool
LLM QUALIFICATION INCOMPLETE: 3/6 measured; the remainder were
  provider/infrastructure non-measurements, not results

Verified

Check Result
bash -n clean
PG-only (no declaration) exit 0, QUALIFIED, 20 PG-backed tests ran
0/6 — real provider 429 condition exit 1, QUALIFIED suppressed, all 6 named
3/6 — stubbed go test exit 1, INCOMPLETE, 3 named
6/6 — stubbed go test exit 0, QUALIFIED

The 0/6 branch was reproduced against the actual failure condition, not simulated. Partial/complete used a throwaway stub standing in for go test — not committed, because the repo has no shell-test convention and adding one would widen this change. Those two branches will be exercised by the next real run once provider capacity allows.

Scope discipline

  • No change to model, prompt, decoder, validator, or any production behaviour.
  • MaxTokens: 4096 untouched; the prompt is not shrunk to fit a quota.
  • APA-55 stays uncommitted.
  • Declaring no scenarios leaves the PG-only path unchanged and independently valid — its verdict line now says so explicitly rather than implying a measurement that never happened.
  • claimops-postgres (shared) never contacted; teardown left no claimops-pg-qual container.

One supporting change: the verbose go test output is now tee'd instead of discarded. A record that is never captured cannot be counted, so accounting requires capture.

Closes APA-57. Blocks the APA-55 causal re-measurement until provider capacity is resolved.

The anti-fake-green gap that let an unmeasured run report success.

Observed while attempting the APA-56 causal re-measurement: every live case
self-skipped on provider 429, so zero model measurements were recorded, and
qualify.sh still printed

    QUALIFIED: full suite green against a genuinely fresh, ephemeral database

with exit 0. The existing anti-fake-green check only proves PG-backed tests
RAN; it cannot see a live matrix, so green-by-skipping was indistinguishable
from a real qualification. A verdict is what a ledger records, so a wrong
headline is worse than no gate at all.

The Go harness was honest throughout — it Skipf'd with "the real-LLM matrix
cannot be measured right now. This is a credential/provider condition, not a
result" — which is the only reason this was caught. The harness is left
exactly as it is; the reporting was what lied.

A run now declares its matrix with QUAL_LLM_SCENARIOS, and then:

    requested > 0 AND measured == 0   -> INFRA_BLOCKED, non-zero exit
    0 < measured < expected           -> INCOMPLETE,     non-zero exit
    measured == expected              -> eligible for a verdict

A scenario counts as measured only if the harness emitted an APA49-EVIDENCE
record for it. A 429, a skip, a discard, or a fatal is never a measurement —
those are precisely the outcomes that used to produce a green verdict with
nothing behind it. An incomplete run names the scenarios it did not measure,
so measured and unmeasured are always reported separately.

Declaring no scenarios leaves the PG-only path unchanged and independently
valid; its verdict line now says so explicitly rather than implying a model
measurement that never happened.

The verbose go test output is now tee'd instead of discarded. A record that
is never captured cannot be counted, so accounting requires capture.

No change to model, prompt, decoder, validator, or any production behaviour.
MaxTokens is untouched, the prompt is untouched, and APA-55 stays uncommitted.

Verified: bash -n clean. PG-only run green, exit 0, 20 PG-backed tests. The
0/N branch reproduced against the real provider 429 condition: exit 1,
QUALIFIED suppressed, all 6 scenarios named. Partial and complete branches
verified with a throwaway stub standing in for `go test` (3/6 -> INCOMPLETE
exit 1; 6/6 -> QUALIFIED exit 0); the stub was not committed, because the repo
has no shell-test convention and adding one would widen this change.
@coderabbitai

coderabbitai Bot commented Oct 6, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: d819a6c7-fcdd-4144-81a0-605ce634fc12
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@Aparnap2 Aparnap2 left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed PR #123 against the anti-fake-green failure mode and the qualification contract. PASS / merge-ready.

Verified from the diff: (1) the gate is opt-in via QUAL_LLM_SCENARIOS, so PG-only qualification remains independently valid; (2) measurement is correctly tied to emitted APA49-EVIDENCE records, so 429/skip/discard/fatal cannot count; (3) 0/expected is INFRA_BLOCKED and partial is INCOMPLETE, both non-zero; (4) the full matrix is required before an LLM qualification can return 0; (5) stdout is captured with tee while pipefail preserves go-test status; (6) the declared scenario names are reported when unmeasured; and (7) no model/prompt/decoder/validator/MaxTokens behavior is changed.

The throwaway-stub verification is a test-coverage limitation, not a merge blocker. A durable shell regression test can be a separate follow-up if needed. CI being UNSTABLE is not caused by this diff based on the current commit status; the only reported check here is CodeRabbit success, and the PR notes that qualify.sh is not invoked by CI.

Decision: merge #123. Keep #122 and APA-55 untouched.

@Aparnap2
Aparnap2 merged commit c879e71 into main Oct 6, 2026
5 checks passed
@Aparnap2
Aparnap2 deleted the fix/apa-57-llm-anti-fake-green-gate branch October 6, 2026 12:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant