Repository navigation
fix(apa-54): state the tool/knob ownership contract in the model prompt - #121
Merged
Merged
Conversation
The prompt named conditional knobs without their owning tools, semantic meaning, requiredness, or copy-source, and omitted hash/payload. The authoritative mapping (tool.go:22-35, checkKnob at tool.go:269-280) is now stated verbatim in the prompt, including the rule that a DocumentID from get_documents is not a valid subject_id and get_evidence does not accept subject_id. Adds a behavioral divergence guard (prompt table == Request.Validate table), the previously-unasserted I4 rejection of get_evidence(subject_id=...), and acceptance pins for valid subject_id on each owning tool. checkKnob, decoder, grounding, loop, fixtures untouched.
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configuration
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Aparnap2
marked this pull request as ready for review
October 6, 2026 06:01
Aparnap2
added a commit
that referenced
this pull request
Oct 6, 2026
… prompt (#122) APA-55 Phase-2 live runs under qwen/qwen3.8-27b rejected 17 of 17 reports at one point: orchestrate: missing_additive[0] unknown external source "policy_number" That was an incorrect denial of the sufficient-evidence control (0/3), and because the additive gate rejects first the grounding/evidence/tenant boundaries the campaign exists to measure were never reached. Root cause, from the read-only audit in APA-56: the prompt closed the "kind" set but never stated which vocabulary each kind's "key" is drawn from, and its only worked example carried an empty array. "policy_number" is a real claim field (xdoc.KeyPolicyNumber, an R1 identifier) paired with the wrong kind. This is a documentation gap in the model-facing contract, the same shape as the knob-ownership gap fixed in #121. The validator is correct and is not weakened: action.go still rejects an unknown external key, and a regression test pins that exact pair so it can never be accepted. Only the prompt and one additive accessor change. The prompt now states the kind -> key mapping for all three kinds, names policy_number explicitly as kind "field", states the carry-forward rule that grounding's additive-superset check requires, and states the validator's real sort tuple (kind, key, detail). A non-empty canonical example was added for each kind and is executed against the real decoder, validator, and grounding in the tests, because an example the contract rejects teaches the model to fail. The external vocabulary is read from invest.ExternalSourceKeys() rather than re-spelled in the prompt, so the two cannot drift; a divergence test fails if they do. Adding a fifth external source without updating the prompt is a test failure, not a silent prompt/model divergence. The action.go ↔ exception.go duplication of the four external keys is deliberately left alone: removing it is a refactor, not this fix, and is recorded as technical debt rather than folded in here. Gates: gofmt clean, go build, go vet, full go test ./... green against a fresh ephemeral PostgreSQL (20 PG-backed tests ran), uv run pytest tests/unit 27 passed, ruff clean. eval/ and tests/ untouched. No real-LLM run in this commit, by design; the causal re-measurement is a separate qualification step.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What was incomplete
The prompt named conditional knobs (cursor, query, subject_id, source_type) without their owning tools, semantic meaning, requiredness, or copy-source, and omitted hash/payload entirely. The model had no authoritative mapping to follow.
Authoritative mapping used
Copied verbatim from tool.go:22-35 (package doc) and checkKnob at tool.go:269-280 (enforced via Request.Validate at tool.go:229-247). checkKnob itself was not modified.
Tests added
Gates
gofmt clean, go build ./..., go vet ./..., go test ./... green, -race green on touched packages, uv run pytest tests/unit 27 passed, ruff check . clean, git diff main --stat -- eval/ tests/ workflows/ empty.