Skip to content

fix(apa-56): state the missing_additive kind/key mapping in the model prompt - #122

Merged
Aparnap2 merged 1 commit into
mainfrom
fix/apa-56-additive-contract
Oct 6, 2026
Merged

Aparnap2 merged 1 commit into
mainfrom
fix/apa-56-additive-contract

Conversation

@Aparnap2

@Aparnap2 Aparnap2 commented Oct 6, 2026

Copy link
Copy Markdown
Owner

What this fixes

APA-55 Phase-2 live runs under qwen/qwen3.8-27b rejected 17 of 17 reports at a single point:

orchestrate: missing_additive[0] unknown external source "policy_number"

That was an incorrect denial of the sufficient-evidence control (0/3), and because the additive gate rejects first, the grounding/evidence/tenant boundaries the campaign exists to measure were never exercised. BOUNDARY_HELD on all 17 runs evidenced the additive gate, not the gate under test.

Root cause (read-only audit, APA-56)

policy_number is not invented — it is a real authoritative field key (xdoc.KeyPolicyNumber, an R1 identifier, the affected field of RulePolicyNumberConflict). The model paired a real key with the wrong kind.

The prompt closed the kind set but never stated which vocabulary each kind's key comes from, and its only worked example carried an empty array:

- A "missing_additive" entry carries "kind", "key", "detail", with "kind" exactly
  one of "required_document", "field", "external"; keep the array sorted.

Grep proved "tpa", the four-key enumeration, MissingExternal, and "external source" appear nowhere in model.go except the incidental tool name get_tpa_case. Zero in-context signal for the key vocabulary.

Classification: missing documentation in the model-facing contract — the same shape as the knob-ownership/subject_id gap fixed in #121. Not a validator defect. Not a model defect.

What is NOT changed

  • action.go validator — still rejects the unknown external key
  • invest/exception.go authoritative vocabulary — unchanged
  • grounding, decoder, invariants, APA-55 acceptance criteria, eval, workflow

What changes

model.go — the prompt now states:

  • the kind → key mapping for all three kinds (required_document / field / external)
  • that policy_number is a claim field, so kind:"field", and is never kind:"external"
  • the carry-forward rule grounding's additive-superset check requires
  • the validator's real sort tuple (kind, key, detail), not just "sorted"
  • a non-empty canonical example per kind

invest/exception.go — one additive read-only accessor, ExternalSourceKeys(), so the prompt is derived from the authority rather than re-spelling it. No vocabulary change; returns a copy.

Tests

apa56_additive_contract_test.go, 6 tests / 8 subtests:

Test Role
PromptStatesTheKindKeyContract the correction is present (4 subtests)
ExternalVocabularyMatchesAuthority divergence test — prompt keys ≡ invest.ExternalSourceKeys()
PolicyNumberClassifiedAsFieldIsAccepted corrected pair is accepted end-to-end
PolicyNumberClassifiedAsExternalIsRejected regression pin on the observed failure
CanonicalNonEmptyExamplesSurviveTheContract every non-empty example is executed against the real decoder/validator/grounding

RED proof: reverting only model.go fails 6 of 8 — including the divergence test and the non-empty-example test. The two validator pins pass before and after by design: they guard against future weakening, they are not the proof for this fix.

The divergence test is the durable part: adding a fifth external source without updating the prompt becomes a test failure rather than a silent prompt/model divergence.

Gates

  • gofmt clean · go build ./... · go vet ./... · go test -count=1 -v -run TestAPA56 8/8 PASS
  • Full go test ./... green against a fresh ephemeral PostgreSQL — QUALIFIED, exit 0, 20 PG-backed tests ran (not self-skipped)
  • uv run pytest tests/unit 27 passed · ruff clean
  • eval/ and tests/ untouched

Deliberately out of scope

action.go:216-221 re-spells the four external keys inline, duplicating exception.go. Removing that is a refactor, not this fix, so it is recorded as technical debt rather than folded in here. The prompt side is single-sourced; the validator side is not yet.

No real-LLM run in this PR, by design. The causal re-measurement (APA-55 Live Evidence under Qwen after this merges) is the next step and must observe the control's false-negative disappear.

Closes APA-56. Blocks APA-55.

🤖 Generated with Claude Code

… prompt

APA-55 Phase-2 live runs under qwen/qwen3.8-27b rejected 17 of 17 reports at
one point:

    orchestrate: missing_additive[0] unknown external source "policy_number"

That was an incorrect denial of the sufficient-evidence control (0/3), and
because the additive gate rejects first the grounding/evidence/tenant
boundaries the campaign exists to measure were never reached.

Root cause, from the read-only audit in APA-56: the prompt closed the "kind"
set but never stated which vocabulary each kind's "key" is drawn from, and
its only worked example carried an empty array. "policy_number" is a real
claim field (xdoc.KeyPolicyNumber, an R1 identifier) paired with the wrong
kind. This is a documentation gap in the model-facing contract, the same
shape as the knob-ownership gap fixed in #121.

The validator is correct and is not weakened: action.go still rejects an
unknown external key, and a regression test pins that exact pair so it can
never be accepted. Only the prompt and one additive accessor change.

The prompt now states the kind -> key mapping for all three kinds, names
policy_number explicitly as kind "field", states the carry-forward rule that
grounding's additive-superset check requires, and states the validator's real
sort tuple (kind, key, detail). A non-empty canonical example was added for
each kind and is executed against the real decoder, validator, and grounding
in the tests, because an example the contract rejects teaches the model to
fail.

The external vocabulary is read from invest.ExternalSourceKeys() rather than
re-spelled in the prompt, so the two cannot drift; a divergence test fails if
they do. Adding a fifth external source without updating the prompt is a test
failure, not a silent prompt/model divergence.

The action.go ↔ exception.go duplication of the four external keys is
deliberately left alone: removing it is a refactor, not this fix, and is
recorded as technical debt rather than folded in here.

Gates: gofmt clean, go build, go vet, full go test ./... green against a
fresh ephemeral PostgreSQL (20 PG-backed tests ran), uv run pytest tests/unit
27 passed, ruff clean. eval/ and tests/ untouched. No real-LLM run in this
commit, by design; the causal re-measurement is a separate qualification
step.
@coderabbitai

coderabbitai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 6f0cd4aa-f36d-4a5a-964e-8ac5d7160e41
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@Aparnap2
Aparnap2 marked this pull request as ready for review October 6, 2026 12:29
@Aparnap2
Aparnap2 merged commit d778ddc into main Oct 6, 2026
5 checks passed
@Aparnap2
Aparnap2 deleted the fix/apa-56-additive-contract branch October 6, 2026 12:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant