Skip to content

fix(apa-54): state the tool/knob ownership contract in the model prompt - #121

Merged
Aparnap2 merged 1 commit into
mainfrom
fix/apa-54-knob-ownership-contract
Oct 6, 2026
Merged

Aparnap2 merged 1 commit into
mainfrom
fix/apa-54-knob-ownership-contract

Conversation

@Aparnap2

@Aparnap2 Aparnap2 commented Oct 6, 2026

Copy link
Copy Markdown
Owner

What was incomplete

The prompt named conditional knobs (cursor, query, subject_id, source_type) without their owning tools, semantic meaning, requiredness, or copy-source, and omitted hash/payload entirely. The model had no authoritative mapping to follow.

Authoritative mapping used

Copied verbatim from tool.go:22-35 (package doc) and checkKnob at tool.go:269-280 (enforced via Request.Validate at tool.go:229-247). checkKnob itself was not modified.

Tests added

  • TestKnobOwnership_GetEvidenceRejectsSubjectID — the previously-unasserted I4 ownership rejection: get_evidence(subject_id=<document_id>) fails closed at the executor gate, backend not invoked.
  • TestKnobOwnership_ValidSubjectIDAccepted — valid subject_id accepted at Request.Validate on all five owning tools; rejected on get_evidence.
  • TestKnobOwnership_PromptMatchesAuthoritativeTable — divergence guard: the prompt's stated knob→tool table must equal the table derived behaviorally from the real Request.Validate, both directions.
  • TestKnobOwnership_PromptStatesTheSubjectIDRule — pins the DocumentID-is-not-a-subject_id rule.
  • TestKnobOwnership_PromptDefinesSubjectIDPerTool — pins per-tool subject_id semantics.
  • TestKnobOwnership_PromptNamesHashAndPayload — pins hash/payload presence.

Gates

gofmt clean, go build ./..., go vet ./..., go test ./... green, -race green on touched packages, uv run pytest tests/unit 27 passed, ruff check . clean, git diff main --stat -- eval/ tests/ workflows/ empty.

The prompt named conditional knobs without their owning tools, semantic
meaning, requiredness, or copy-source, and omitted hash/payload. The
authoritative mapping (tool.go:22-35, checkKnob at tool.go:269-280) is
now stated verbatim in the prompt, including the rule that a DocumentID
from get_documents is not a valid subject_id and get_evidence does not
accept subject_id. Adds a behavioral divergence guard (prompt table ==
Request.Validate table), the previously-unasserted I4 rejection of
get_evidence(subject_id=...), and acceptance pins for valid subject_id
on each owning tool. checkKnob, decoder, grounding, loop, fixtures
untouched.
@coderabbitai

coderabbitai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 3aece83b-8809-4530-b8ad-3648ab8b3995
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@Aparnap2
Aparnap2 marked this pull request as ready for review October 6, 2026 06:01
@Aparnap2
Aparnap2 merged commit f2d8466 into main Oct 6, 2026
5 checks passed
@Aparnap2
Aparnap2 deleted the fix/apa-54-knob-ownership-contract branch October 6, 2026 06:01
Aparnap2 added a commit that referenced this pull request Oct 6, 2026
… prompt (#122)

APA-55 Phase-2 live runs under qwen/qwen3.8-27b rejected 17 of 17 reports at
one point:

    orchestrate: missing_additive[0] unknown external source "policy_number"

That was an incorrect denial of the sufficient-evidence control (0/3), and
because the additive gate rejects first the grounding/evidence/tenant
boundaries the campaign exists to measure were never reached.

Root cause, from the read-only audit in APA-56: the prompt closed the "kind"
set but never stated which vocabulary each kind's "key" is drawn from, and
its only worked example carried an empty array. "policy_number" is a real
claim field (xdoc.KeyPolicyNumber, an R1 identifier) paired with the wrong
kind. This is a documentation gap in the model-facing contract, the same
shape as the knob-ownership gap fixed in #121.

The validator is correct and is not weakened: action.go still rejects an
unknown external key, and a regression test pins that exact pair so it can
never be accepted. Only the prompt and one additive accessor change.

The prompt now states the kind -> key mapping for all three kinds, names
policy_number explicitly as kind "field", states the carry-forward rule that
grounding's additive-superset check requires, and states the validator's real
sort tuple (kind, key, detail). A non-empty canonical example was added for
each kind and is executed against the real decoder, validator, and grounding
in the tests, because an example the contract rejects teaches the model to
fail.

The external vocabulary is read from invest.ExternalSourceKeys() rather than
re-spelled in the prompt, so the two cannot drift; a divergence test fails if
they do. Adding a fifth external source without updating the prompt is a test
failure, not a silent prompt/model divergence.

The action.go ↔ exception.go duplication of the four external keys is
deliberately left alone: removing it is a refactor, not this fix, and is
recorded as technical debt rather than folded in here.

Gates: gofmt clean, go build, go vet, full go test ./... green against a
fresh ephemeral PostgreSQL (20 PG-backed tests ran), uv run pytest tests/unit
27 passed, ruff clean. eval/ and tests/ untouched. No real-LLM run in this
commit, by design; the causal re-measurement is a separate qualification
step.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant