Repository navigation
fix(apa-56): state the missing_additive kind/key mapping in the model prompt - #122
Merged
Merged
Conversation
… prompt
APA-55 Phase-2 live runs under qwen/qwen3.8-27b rejected 17 of 17 reports at
one point:
orchestrate: missing_additive[0] unknown external source "policy_number"
That was an incorrect denial of the sufficient-evidence control (0/3), and
because the additive gate rejects first the grounding/evidence/tenant
boundaries the campaign exists to measure were never reached.
Root cause, from the read-only audit in APA-56: the prompt closed the "kind"
set but never stated which vocabulary each kind's "key" is drawn from, and
its only worked example carried an empty array. "policy_number" is a real
claim field (xdoc.KeyPolicyNumber, an R1 identifier) paired with the wrong
kind. This is a documentation gap in the model-facing contract, the same
shape as the knob-ownership gap fixed in #121.
The validator is correct and is not weakened: action.go still rejects an
unknown external key, and a regression test pins that exact pair so it can
never be accepted. Only the prompt and one additive accessor change.
The prompt now states the kind -> key mapping for all three kinds, names
policy_number explicitly as kind "field", states the carry-forward rule that
grounding's additive-superset check requires, and states the validator's real
sort tuple (kind, key, detail). A non-empty canonical example was added for
each kind and is executed against the real decoder, validator, and grounding
in the tests, because an example the contract rejects teaches the model to
fail.
The external vocabulary is read from invest.ExternalSourceKeys() rather than
re-spelled in the prompt, so the two cannot drift; a divergence test fails if
they do. Adding a fifth external source without updating the prompt is a test
failure, not a silent prompt/model divergence.
The action.go ↔ exception.go duplication of the four external keys is
deliberately left alone: removing it is a refactor, not this fix, and is
recorded as technical debt rather than folded in here.
Gates: gofmt clean, go build, go vet, full go test ./... green against a
fresh ephemeral PostgreSQL (20 PG-backed tests ran), uv run pytest tests/unit
27 passed, ruff clean. eval/ and tests/ untouched. No real-LLM run in this
commit, by design; the causal re-measurement is a separate qualification
step.
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configuration
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this fixes
APA-55 Phase-2 live runs under
qwen/qwen3.8-27brejected 17 of 17 reports at a single point:That was an incorrect denial of the sufficient-evidence control (0/3), and because the additive gate rejects first, the grounding/evidence/tenant boundaries the campaign exists to measure were never exercised.
BOUNDARY_HELDon all 17 runs evidenced the additive gate, not the gate under test.Root cause (read-only audit, APA-56)
policy_numberis not invented — it is a real authoritative field key (xdoc.KeyPolicyNumber, an R1 identifier, the affected field ofRulePolicyNumberConflict). The model paired a real key with the wrong kind.The prompt closed the
kindset but never stated which vocabulary each kind'skeycomes from, and its only worked example carried an empty array:Grep proved
"tpa", the four-key enumeration,MissingExternal, and "external source" appear nowhere inmodel.goexcept the incidental tool nameget_tpa_case. Zero in-context signal for the key vocabulary.Classification: missing documentation in the model-facing contract — the same shape as the knob-ownership/
subject_idgap fixed in #121. Not a validator defect. Not a model defect.What is NOT changed
action.govalidator — still rejects the unknown external keyinvest/exception.goauthoritative vocabulary — unchangedWhat changes
model.go— the prompt now states:kind→keymapping for all three kinds (required_document/field/external)policy_numberis a claim field, sokind:"field", and is neverkind:"external"(kind, key, detail), not just "sorted"invest/exception.go— one additive read-only accessor,ExternalSourceKeys(), so the prompt is derived from the authority rather than re-spelling it. No vocabulary change; returns a copy.Tests
apa56_additive_contract_test.go, 6 tests / 8 subtests:PromptStatesTheKindKeyContractExternalVocabularyMatchesAuthorityinvest.ExternalSourceKeys()PolicyNumberClassifiedAsFieldIsAcceptedPolicyNumberClassifiedAsExternalIsRejectedCanonicalNonEmptyExamplesSurviveTheContractRED proof: reverting only
model.gofails 6 of 8 — including the divergence test and the non-empty-example test. The two validator pins pass before and after by design: they guard against future weakening, they are not the proof for this fix.The divergence test is the durable part: adding a fifth external source without updating the prompt becomes a test failure rather than a silent prompt/model divergence.
Gates
gofmtclean ·go build ./...·go vet ./...·go test -count=1 -v -run TestAPA568/8 PASSgo test ./...green against a fresh ephemeral PostgreSQL —QUALIFIED, exit 0, 20 PG-backed tests ran (not self-skipped)uv run pytest tests/unit27 passed ·ruffcleaneval/andtests/untouchedDeliberately out of scope
action.go:216-221re-spells the four external keys inline, duplicatingexception.go. Removing that is a refactor, not this fix, so it is recorded as technical debt rather than folded in here. The prompt side is single-sourced; the validator side is not yet.No real-LLM run in this PR, by design. The causal re-measurement (APA-55 Live Evidence under Qwen after this merges) is the next step and must observe the control's false-negative disappear.
Closes APA-56. Blocks APA-55.
🤖 Generated with Claude Code