Skip to content

Adversarial test suite for AUTO mode: standing grants, quota boundaries, the standing marker #104

Description

@hsliuustc0106

#67 made AUTO the default mode: standing pre-grants derive content-bound single-use capabilities with no prompt (#48 decision 5). The gated path's adversarial thickness was built when gated was the default; AUTO's blast radius is now the live one. This issue pins AUTO's boundary semantics with named adversarial tests (mostly against PermissionCenter.authorize_auto, WriteQuota, and the write flow's AUTO path).

Acceptance (each row a named offline test)

  • Exact-match authority: a standing grant for (action, target, scope, task) authorizes exactly that tuple; any of target/other-action/other-scope/other-task mismatches return None with zero capability rows created
  • Expiry & revocation: an expired or revoked standing grant derives nothing; a capability derived before revocation but executing after a scope change is stopped (supersession / pending-capability join), never sent
  • The standing marker is not a wildcard: STANDING_CONTENT_HASH on a grant can never satisfy a payload check — a capability whose content hash equals the standing marker sends nothing (the writer's wire hash check fails closed); approve() never issues a capability carrying the standing marker
  • Mode discipline: authorize_auto outside AUTO mode raises WriteForbidden; readonly mode never reaches any authorize path
  • Quota day boundaries (FakeClock across midnight): check before the boundary and consume after lands in the new day; exhaustion today blocks both proposal and execution today, and the next day recovers without restart
  • Quota exhaustion degrades, never errors: an exhausted budget mid-execution records write-failed: quota exhausted, burns the capability (no retry), and the run outcome stays healthy
  • Silent-failure visibility: with no matching standing grant, a would-be write is skipped and recorded exactly once (the auto-seen scan dedups) — a broken authorize path is diagnosable from the log, not silent
  • Full suite + demo green

Out of scope

Changing any semantics — this issue adds tests that pin current intended behavior; if a row FAILS against current behavior, that's a finding to triage in a comment first.

Activity

added a commit that references this issue on Oct 4, 2026

hsliuustc0106 commented on Oct 4, 2026

@hsliuustc0106
ContributorAuthor

Implemented and merged with #106 (CI green). All rows verified; one expectation tightened (action mismatch raises WriteForbidden from the mode gate — stricter than the proposed None, kept and documented).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions