Skip to content

feat(core): enforce controlled benchmark acceptance thresholds - #237

Open
seonghobae wants to merge 59 commits into
mainfrom
feat/controlled-benchmark-threshold-contract
Open

seonghobae wants to merge 59 commits into
mainfrom
feat/controlled-benchmark-threshold-contract

Conversation

@seonghobae

@seonghobae seonghobae commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

Scope

Controlled benchmark release-readiness and threshold contract on protected main@87c4daa1830bac5a5228b6036752ad5633232085. This lane owns OriginWeave benchmark evidence admission and acceptance semantics only; it does not duplicate release/signing authority or sibling-owner logic.

Current exact head is 4d175467c550c969d1ad22e51016473c0a4da034, open / Ready / mergeable, directly based on protected main.

Latest executable RED and repair

Predecessor exact ea92c326e2dc4e3daa869aff1266c10b05453e7d received native CI and exposed two genuine defects: rustfmt rejected the hostile control-character fixture layout, and production coverage missed the ControlledBenchmarkSuiteError::ControlCharacterRunContext Display arm. Ordinary-forward test-only commit 4d175467... repaired both without changing production semantics, thresholds, workflow policy, or release authority.

Exact-current repository/security evidence

On unchanged exact 4d175467...:

  • Manifest V3 Compatibility 35689677725: SUCCESS;
  • CI 35689677783: SUCCESS;
    • Production coverage 106623791231: SUCCESS, including exact production coverage enforcement;
    • Rust contracts 106623791494: SUCCESS, including repository contracts, rustfmt, locked tests, strict Clippy and API docs;
  • SAST Semgrep 35689677835: SUCCESS;
  • Security Scan 35689677811: SUCCESS;
  • CodeQL PR 35689677975: FAILURE, with the current causal boundary now isolated in the canonical .github producer rather than OriginWeave source/SARIF.

CodeQL owner-path state — current

The prior body is superseded: the central language jobs are no longer queued.

Leaf Detect 106623791602 succeeded. Compatibility jobs 106663165098 (Python), 106663165121 (JavaScript/TypeScript), and 106663165232 (Actions) fail-closed while awaiting an authenticated terminal dispatch verdict; coordinator 106704274687 later revalidated the same head/base and successfully dispatched the exact scan.

Canonical .github producer 35729253661 has now fully executed its current jobs:

  • validate-dispatch 106750442061: SUCCESS, including live target-PR metadata binding;
  • Actions 106824378689: CodeQL initialization, actual analysis, and Medium+ SARIF gate all SUCCESS, then FAILURE at Verify GHAS base/head CodeQL configuration identity;
  • Python 106824378827: same sequence, failing only at GHAS configuration identity verification after analysis/SARIF success;
  • JavaScript/TypeScript 106824378910: same sequence;
  • settle exact required run 106864104796: App-token exchange succeeds, then FAILURE at Settle exact CodeQL required run.

The exact current blocker is therefore canonical GHAS base/head configuration-identity verification plus exact cross-repository required-run settlement, not scanner admission and not an established OriginWeave source/SARIF defect. The unchanged-head canary is recorded in ContextualWisdomLab/.github#1929 comment 5803453677. Keep the leaf fail-closed; do not blind-rerun, synthesize statuses, copy the central workflow into OriginWeave, broaden credentials, or manufacture a no-op wake.

Review authority — current

Fresh formal review inventory on this exact head contains both:

  • Noema PRR_kwDOTulPlM8AAAABOoVPNA: APPROVED, explicitly bound to 4d175467...; and
  • OpenCode PRR_kwDOTulPlM8AAAABOtL90w: CHANGES_REQUESTED, also bound to 4d175467..., because central OpenCode coverage-evidence run 35714835267 failed before model review at Measure test and docstring evidence.

The OpenCode dispatch metadata job succeeded, but central coverage-evidence job 106752951786 failed at that measurement step. Its later opencode-review job 106825479498 skipped the model pool, published the formal fail-closed COVERAGE_BLOCKED review outcome, and failed its wake/status publication tail. This remains a governance blocker even though native OriginWeave Production coverage is GREEN; no attempt is made here to dismiss or overwrite the review. Fresh review-thread inventory is empty.

The canonical autonomous repair owner for Required OpenCode coverage-evidence failures is ContextualWisdomLab/.github#2169 / successor #2170. Fresh owner-path inspection found #2170's live GitHub head has advanced to 422ba67ce3eba06632266a0f6452754b63c50052, while its PR body still names an older generation. Exact 422ba67c... is currently not accepted: Security Scan succeeds, but Agent Review Runtime Quality fails at Verify exact-head path policy and syntax, Semgrep generates/uploads SARIF and fails its Medium+ gate, Bandit generates/uploads SARIF and fails its MEDIUM+ gate, and CodeQL is cancelled. This newer state plus the unchanged #237 coverage canary is handed to .github#2169 in comment 5805625784. Do not duplicate the central scheduler repair in OriginWeave or manufacture a leaf wake; owner #2170 must read/adopt its intervening delta, clear those exact-head findings, land normally, then this unchanged head must be reevaluated.

Dependent stack

#322 remains based on predecessor #237 exact ea92c326...; current #322 exact a4c8ceaf67a075ef483334802aacfc54cf502068 is therefore one parent repair behind and must ordinary/non-force adopt/adapt the accepted parent generation after #237 integration, preserving its broader diagnostic contract. #324 remains downstream of #322; its existing leaf GREEN becomes predecessor evidence after any restack.

Integration remains ancestor-first: #237 acceptance and normal protected integration → #322 ordinary/non-force adoption plus fresh exact-head acceptance/integration → #324 ordinary/non-force adoption plus fresh exact-head acceptance/integration.

Acceptance

Repository/MV3/SAST/Security are GREEN on this exact head, and Noema has an exact-head approval with zero review threads. The PR is nevertheless not merge-accepted: required CodeQL remains terminal RED at the canonical GHAS identity/settlement boundary, and the current exact-head OpenCode review is CHANGES_REQUESTED following its central coverage-evidence failure. Both owner paths must reach authentic terminal acceptance on the unchanged exact head, then the effective ruleset must be re-read before normal protected-main integration.

No self-approval, review dismissal, bypass, force push, destructive rebase, workflow/ruleset/secret mutation, gate weakening, blind rerun, source-neutral wake, protected-main merge, tag, package, publish or release is authorized by this state.

@coderabbitai

coderabbitai Bot commented Aug 27, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: b6b4b47f-5802-47c5-af88-ed92347ba762

📥 Commits

Reviewing files that changed from the base of the PR and between ea92c32 and 4d17546.

📒 Files selected for processing (1)
  • crates/originweave-core/tests/controlled_benchmark_run_context.rs

Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

originweave-core에 버전 관리형 controlled benchmark registry와 100회 trial 기반 evidence 평가를 추가했다. 지원 프로필, 실행 context, malformed evidence를 검증하고 case 및 suite outcome을 계산한다. 관련 통합 테스트와 문서를 추가했다.

Changes

결정론적 benchmark 평가

Layer / File(s) Summary
Benchmark 계약과 evidence 모델
crates/originweave-core/src/controlled_benchmark.rs, crates/originweave-core/src/root.rs
고정 trial 수, 20개 case identity, 지원 프로필, 실행 context, raw trial 및 aggregate evidence 구조와 오류 타입을 추가했다.
Trial 집계와 case 판정
crates/originweave-core/src/controlled_benchmark.rs, crates/originweave-core/tests/controlled_benchmark.rs, crates/originweave-core/tests/controlled_benchmark_trial_aggregation.rs, crates/originweave-core/tests/controlled_benchmark_unauthorized_side_effect_count.rs, crates/originweave-core/tests/controlled_benchmark_zero_trial_side_effect.rs
trial ordinal, counter, side effect overflow를 검증한다. 모든 조건을 충족한 100회는 Passed, 알려진 실패는 Failed, 부족한 trial은 Inconclusive로 판정한다.
Run context와 suite 판정
crates/originweave-core/src/controlled_benchmark.rs, crates/originweave-core/tests/controlled_benchmark*.rs
실행 context의 canonical 형식과 byte-for-byte 일치를 검증한다. registry version, 지원 프로필, required case, 중복 및 conditional case를 검증한 뒤 suite outcome을 반환한다.
아키텍처와 변경 기록
ARCHITECTURE.md, CHANGELOG.md
controlled benchmark registry와 raw evidence 평가 범위, 100회 trial 규칙, required case 및 fail-closed 규칙을 문서화했다.

Priority: ⬇️ Low

Estimated code review effort: 4 (Complex) | ~60 minutes

Change: Feature

Sequence Diagram(s)

sequenceDiagram
  participant Caller
  participant RunEvaluator
  participant SuiteEvaluator
  participant CaseEvaluator
  Caller->>RunEvaluator: 실행 context와 benchmark evidence 전달
  RunEvaluator->>SuiteEvaluator: 검증된 실행 정보 전달
  SuiteEvaluator->>CaseEvaluator: required case evidence 전달
  CaseEvaluator-->>SuiteEvaluator: case outcome 반환
  SuiteEvaluator-->>Caller: suite outcome 반환
Loading

Merge Risk: ⚪ Minimal · up to 4d175

This PR adds bounded, fail-closed evaluation of controlled benchmark evidence without replacing benchmark execution or the product release gate. No merge-blocking risk remains; it is mergeable after normal checks.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed Docstring coverage is 84.21% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 57 functions across 10 files.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed 제목은 controlled benchmark의 acceptance threshold 강제라는 PR의 주요 변경을 정확하고 간결하게 설명합니다.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@seonghobae

Copy link
Copy Markdown
Contributor Author

Current HEAD: 5d2e67b; base: 542ca1e.

Added fail-closed validation for impossible unauthorized_side_effects counts through the shared aggregate-counter validator, with a regression test. Updated ARCHITECTURE.md and CHANGELOG.md for the active controlled-benchmark boundary.

Local evidence on this exact HEAD: focused controlled/release acceptance tests passed; cargo test --workspace --all-features passed on Rust 1.97.1; cargo clippy -D warnings, rustdoc -D warnings, cargo fmt --check, Python unittest (153), compileall, git diff --check, and production functions/lines/regions/branches coverage (100%) passed. Hosted checks are newly queued; no predecessor evidence is being reused. PR remains draft and no merge is claimed.

@seonghobae

Copy link
Copy Markdown
Contributor Author

Exact-head maintenance audit for 53f816b3d4bfee995135ff5fb7209e251563dbd6 against protected main 542ca1e9c0a863595b8b6697790005d2471f5413:

  • Corrected the PR description's stale exact-head evidence: it now names the live head and reports only current checks.
  • Current coverage-evidence, Rust contracts, Production coverage, security/SAST, Python, Noema, and repository checks are successful. opencode-review is queued and is not treated as passing.
  • Current inline review-thread inventory is empty; no formal independent approval is present.
  • This remains a Draft partial implementation of [Product Gap] Establish a release-grade web-agent benchmark and commercial acceptance gate #203: no real browser runner, durable benchmark evidence pipeline, or full five-suite commercial gate is claimed. No merge or protected-gate bypass was made.

@seonghobae

Copy link
Copy Markdown
Contributor Author

Exact-head blocker update for 53f816b3d4bfee995135ff5fb7209e251563dbd6 against protected main@542ca1e9c0a863595b8b6697790005d2471f5413:

  • Local/hosted Rust contracts, Production coverage, coverage-evidence, security scans, Python analysis, Noema review, and workflow checks are successful for this exact head.
  • opencode-review is a completed failure: job step Fail closed without a current-head OpenCode verdict failed, so no authenticated current-head OpenCode verdict exists.
  • strix is still in_progress; it is not passing evidence while running.
  • Current review inventory and inline review threads are empty, but the PR remains Draft and has no qualifying independent approval.
  • The PR body was updated to reflect the current failure/in-progress states; no predecessor evidence or merge was used.

@seonghobae

Copy link
Copy Markdown
Contributor Author

Exact-head blocker RCA update for 53f816b3d4bfee995135ff5fb7209e251563dbd6:

  • strix completed with failure in job 98772183805 / run 33147195777, after bounded retries. The current log records STRIX_PROVIDER_UNAVAILABLE; the contextual-orchestrator gateway returned HTTP 400 invalid_tools because a tool function description exceeded the provider limit of 1024 characters, then the workflow failed closed.
  • This is an organization/provider backend failure and supplies no authoritative vulnerability verdict; there is no safe OriginWeave product-code change that can repair that external gateway contract.
  • opencode-review remains a completed failure for the same exact head, while coverage-evidence, Rust contracts, Production coverage, and the other repository gates remain successful.
  • The PR body now records both completed failures and does not promote them as passing evidence. The PR remains Draft with no qualifying independent approval; no bypass or merge was attempted.

@seonghobae

Copy link
Copy Markdown
Contributor Author

Current-main reconciliation is now pushed at exact head 60b7ef15f645d50853f26cfdafcf90df81184b73 (signed merge commit; non-force fast-forward). The prior Strix failure was not a source finding: check 99062402046 reports STRIX_PROVIDER_UNAVAILABLE; its other annotation is only exit code 1. On the reconciled tree, focused controlled-benchmark tests passed (32), all Python repository contracts passed (152), workspace fmt/Clippy/all-target all-feature tests/rustdoc passed, and the pinned nightly LLVM report verified 100% functions, lines, regions, and branches. This is local evidence only: keep Draft until the new exact-head hosted checks and current review are terminal; do not transfer predecessor approvals or statuses.

@seonghobae

Copy link
Copy Markdown
Contributor Author

Correction to the preceding evidence comment: the exact pushed head is 60b7ef1534c44b0cb2b15d6929b184a8e0c7a0e1. The abbreviated 60b7ef15 was correct; the typed long SHA in that comment was not. All local verification described there ran on the actual exact head above.

Copy link
Copy Markdown
Contributor Author

@coderabbitai review

Please review exact head d5ddf5784d5a81231f3e43965237a0f604cf44f1, focusing on ControlledBenchmarkRunContext admission. Verify that embedded C0/C1 control characters cannot enter benchmark evidence, visible Unicode remains accepted, error/source/rustdoc coverage is exhaustive, and the change does not widen this PR into #199/#201 signing or persistence authority. Also check that the normal non-force protected-main adoption preserved the effective product delta without .github/** ownership.

@seonghobae
seonghobae marked this pull request as ready for review September 14, 2026 21:37
@coderabbitai

coderabbitai Bot commented Sep 14, 2026

Copy link
Copy Markdown

Tip

For best results, initiate chat on the files or code changes.

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo '== repository state and requested commits =='
git status --short
git rev-parse HEAD
git cat-file -t d5ddf5784d5a81231f3e43965237a0f604cf44f1
git cat-file -t 87c4daa1830bac5a5228b6036752ad5633232085
git merge-base --is-ancestor 87c4daa1830bac5a5228b6036752ad5633232085 d5ddf5784d5a81231f3e43965237a0f604cf44f1 && echo 'base_is_ancestor=yes'

echo '== effective changed paths: base...head =='
git diff --name-status 87c4daa1830bac5a5228b6036752ad5633232085...d5ddf5784d5a81231f3e43965237a0f604cf44f1
echo '== effective .github paths =='
git diff --name-status 87c4daa1830bac5a5228b6036752ad5633232085...d5ddf5784d5a81231f3e43965237a0f604cf44f1 -- .github || true

echo '== PR metadata =='
gh pr view 237 --repo ContextualWisdomLab/OriginWeave --json number,state,isDraft,mergeable,headRefOid,baseRefOid,files --jq '{number,state,isDraft,mergeable,headRefOid,baseRefOid,files:[.files[].path]}'

echo '== candidate controlled benchmark files =='
fd -t f . | rg -i 'controlled|benchmark|acceptance|architecture|changelog'

echo '== Rust symbol map =='
files=$(fd -t f -e rs . | xargs -r rg -l 'ControlledBenchmarkRunContext|ControlCharacterRunContext' || true)
if [ -n "$files" ]; then
  printf '%s\n' "$files"
  printf '%s\n' "$files" | xargs -r ast-grep outline --items all --type struct,enum,impl,function
fi

echo '== symbol occurrences with context =='
rg -n -C 5 'ControlledBenchmarkRunContext|ControlCharacterRunContext|is_control|control character' -g '*.rs' . || true

echo '== effective patch =='
git diff --find-renames --find-copies 87c4daa1830bac5a5228b6036752ad5633232085...d5ddf5784d5a81231f3e43965237a0f604cf44f1 -- '*.rs' 'ARCHITECTURE.md' 'CHANGELOG.md'

Length of output: 5124


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo '== source outline =='
ast-grep outline crates/originweave-core/src/controlled_benchmark.rs --items all --type struct,enum,function,method,trait

echo '== run-context test outline =='
ast-grep outline crates/originweave-core/tests/controlled_benchmark_run_context.rs --items all --type function

echo '== exact control-character repair patch =='
git show --format=fuller --find-renames --find-copies d5ddf5784d5a81231f3e43965237a0f604cf44f1 -- \
  crates/originweave-core/src/controlled_benchmark.rs \
  crates/originweave-core/tests/controlled_benchmark_run_context.rs

echo '== controlled benchmark source =='
cat -n crates/originweave-core/src/controlled_benchmark.rs

echo '== run-context tests =='
cat -n crates/originweave-core/tests/controlled_benchmark_run_context.rs

echo '== public exports =='
cat -n crates/originweave-core/src/root.rs

echo '== signing/persistence/serialization references in effective changes =='
git diff -U5 87c4daa1830bac5a5228b6036752ad5633232085...d5ddf5784d5a81231f3e43965237a0f604cf44f1 -- \
  ARCHITECTURE.md CHANGELOG.md crates/originweave-core \
  | rg -n -i -C 3 'sign|signature|persist|durable|retain|artifact|hash|serde|serializ|store|database' || true

echo '== all constructors and aggregation admissions =='
rg -n -C 4 'ControlledBenchmarkRunContext\s*\{|run_context:|run_context\)|run_context,' \
  crates/originweave-core/src crates/originweave-core/tests -g '*.rs'

Length of output: 1823


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo '== source outline =='
ast-grep outline crates/originweave-core/src/controlled_benchmark.rs --items all

echo '== run-context test outline =='
ast-grep outline crates/originweave-core/tests/controlled_benchmark_run_context.rs --items all

echo '== exact control-character repair patch =='
git show --format=fuller --find-renames --find-copies d5ddf5784d5a81231f3e43965237a0f604cf44f1 -- \
  crates/originweave-core/src/controlled_benchmark.rs \
  crates/originweave-core/tests/controlled_benchmark_run_context.rs

echo '== controlled benchmark source =='
cat -n crates/originweave-core/src/controlled_benchmark.rs

echo '== run-context tests =='
cat -n crates/originweave-core/tests/controlled_benchmark_run_context.rs

echo '== public exports =='
cat -n crates/originweave-core/src/root.rs

echo '== authority-related effective changes =='
git diff -U5 87c4daa1830bac5a5228b6036752ad5633232085...d5ddf5784d5a81231f3e43965237a0f604cf44f1 -- \
  ARCHITECTURE.md CHANGELOG.md crates/originweave-core \
  | rg -n -i -C 3 'sign|signature|persist|durable|retain|artifact|hash|serde|serializ|store|database' || true

echo '== run-context construction and aggregation use =='
rg -n -C 4 'ControlledBenchmarkRunContext\s*\{|run_context:|run_context\)|run_context,' \
  crates/originweave-core/src crates/originweave-core/tests -g '*.rs' || true

Length of output: 59160


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo '== error-source and display coverage in controlled benchmark tests =='
rg -n -C 5 '\.source\(\)|std::error::Error|ControlCharacterRunContext|InvalidTrialEvidence|to_string\(\)' \
  crates/originweave-core/tests -g 'controlled_benchmark*.rs'

echo '== controlled module documentation references and public API uses =='
rg -n -C 3 'evaluate_controlled_benchmark_suite_for_run|ControlledBenchmarkRunContext|ControlledBenchmarkSuiteError' \
  . -g '*.rs' -g '*.md' -g '!target/**'

echo '== effective diff check and summary =='
git diff --check 87c4daa1830bac5a5228b6036752ad5633232085...d5ddf5784d5a81231f3e43965237a0f604cf44f1
git diff --stat 87c4daa1830bac5a5228b6036752ad5633232085...d5ddf5784d5a81231f3e43965237a0f604cf44f1
git diff --numstat 87c4daa1830bac5a5228b6036752ad5633232085...d5ddf5784d5a81231f3e43965237a0f604cf44f1

echo '== persistence/signing implementation constructs in changed Rust =='
git diff --unified=0 87c4daa1830bac5a5228b6036752ad5633232085...d5ddf5784d5a81231f3e43965237a0f604cf44f1 -- '*.rs' \
  | rg -n '^\+.*\b(fn|struct|enum|impl|use|mod|pub)\b|^\+.*(write|file|path|serialize|deserialize|sign|verify|hash|store|database|http|client)' || true

echo '== changed files at repair commit only =='
git diff-tree --no-commit-id --name-status -r d5ddf5784d5a81231f3e43965237a0f604cf44f1

Length of output: 58323


@seonghobae Review complete for exact head d5ddf5784d5a81231f3e43965237a0f604cf44f1. I found no correctness issue in the requested scope.

  • evaluate_controlled_benchmark_suite_for_run validates all nine fields in both the expected and observed ControlledBenchmarkRunContext values before it compares values or evaluates suite evidence.
  • validate_run_context_field rejects any character accepted by char::is_control. This rejects embedded C0 controls, including LF and NUL, and C1 controls, including U+0085.
  • The test at crates/originweave-core/tests/controlled_benchmark_run_context.rs covers embedded LF, NUL, and U+0085. It expects ControlledBenchmarkSuiteError::ControlCharacterRunContext.
  • The same test accepts the visible Unicode value 결정적-ブラウザ-oráculo-v1. The implementation does not apply an ASCII-only restriction.
  • ControlledBenchmarkSuiteError::ControlCharacterRunContext has variant and field rustdoc, a Display branch, an explicit Error::source() None branch, and public API rustdoc that describes this error condition.
  • The effective diff from protected main 87c4daa1830bac5a5228b6036752ad5633232085 contains no .github/** path. The protected-main commit is an ancestor of the reviewed head.
  • The changed Rust code only evaluates in-memory evidence. It does not implement signing, signature verification, persistence, serialization, artifact storage, or network transport. The documentation explicitly leaves authentication and durable signed evidence to the benchmark runner and evidence pipeline.

I performed static repository inspection only. I did not run repository code or build commands.

You are interacting with an AI system.

@seonghobae
seonghobae marked this pull request as draft September 14, 2026 21:50

Copy link
Copy Markdown
Contributor Author

@coderabbitai review

Please review exact head ea92c326e2dc4e3daa869aff1266c10b05453e7d against protected base 87c4daa1830bac5a5228b6036752ad5633232085, with specific attention to the prior Docstring Coverage warning and whether every production function/method in crates/originweave-core/src/controlled_benchmark.rs now carries decision-relevant rustdoc without changing benchmark semantics. Do not treat hosted CI queue/cancellation as GREEN.

@coderabbitai

coderabbitai Bot commented Sep 14, 2026

Copy link
Copy Markdown

Tip

For best results, initiate chat on the files or code changes.

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

base='87c4daa1830bac5a5228b6036752ad5633232085'
head='ea92c326e2dc4e3daa869aff1266c10b05453e7d'
repo='ContextualWisdomLab/OriginWeave'

echo '== exact commit relationship =='
git rev-parse HEAD
git cat-file -t "$base"
git cat-file -t "$head"
git merge-base --is-ancestor "$base" "$head" && echo 'base_is_ancestor=yes'
git diff --check "$base...$head"

echo '== changed paths =='
git diff --name-status "$base...$head"

echo '== controlled benchmark symbol inventory =='
ast-grep outline crates/originweave-core/src/controlled_benchmark.rs --items all --type function,method,struct,enum,impl

echo '== exact effective patch for controlled benchmark =='
git diff --find-renames --find-copies --unified=80 "$base...$head" -- crates/originweave-core/src/controlled_benchmark.rs

echo '== current source with line numbers =='
cat -n crates/originweave-core/src/controlled_benchmark.rs

echo '== rustdoc lint configuration and doc coverage references =='
rg -n -C 3 'missing_docs|Docstring Coverage|docstring|rustdoc' Cargo.toml crates .github 2>/dev/null || true

echo '== hosted check states for requested head =='
gh api "repos/$repo/commits/$head/check-runs?per_page=100" --jq '.check_runs[] | [.name,.status,.conclusion,.details_url] | `@tsv`' || true
gh pr view 237 --repo "$repo" --json headRefOid,baseRefOid,isDraft,statusCheckRollup --jq '{headRefOid,baseRefOid,isDraft,statusCheckRollup:[.statusCheckRollup[]|{name:.name,status:.status,conclusion:.conclusion}]}' || true

Length of output: 2732


🏁 Script executed:

#!/bin/bash
set -u

base='87c4daa1830bac5a5228b6036752ad5633232085'
head='ea92c326e2dc4e3daa869aff1266c10b05453e7d'
repo='ContextualWisdomLab/OriginWeave'

echo '== source outline =='
ast-grep outline crates/originweave-core/src/controlled_benchmark.rs --items all || true

echo '== all function and method declaration locations =='
rg -n -C 8 '^\s*(pub\s+)?(const\s+)?fn\s+[A-Za-z_][A-Za-z0-9_]*\s*\(' crates/originweave-core/src/controlled_benchmark.rs || true

echo '== source =='
cat -n crates/originweave-core/src/controlled_benchmark.rs || true

echo '== requested-head commit and parent patch =='
git show --no-ext-diff --format=fuller --stat "$head" || true
git diff --no-ext-diff --find-renames --find-copies --unified=30 "$head^" "$head" -- crates/originweave-core/src/controlled_benchmark.rs || true

echo '== effective function-bearing lines in base-head diff =='
git diff --no-ext-diff --unified=3 "$base...$head" -- crates/originweave-core/src/controlled_benchmark.rs | rg -n -C 4 '^\+.*\bfn\s+|^\+.*///|^\+.*//!' || true

echo '== rustdoc lint configuration and documentation coverage references =='
rg -n -C 3 'missing_docs|Docstring Coverage|docstring|rustdoc' Cargo.toml crates .github 2>/dev/null || true

echo '== hosted checks for requested head =='
gh api "repos/$repo/commits/$head/check-runs?per_page=100" --jq '.check_runs[] | [.name,.status,(.conclusion // "null"),.details_url] | `@tsv`' || true

echo '== pull request rollup =='
gh pr view 237 --repo "$repo" --json headRefOid,baseRefOid,isDraft,statusCheckRollup --jq '{headRefOid,baseRefOid,isDraft,statusCheckRollup:[.statusCheckRollup[]|{name:.name,status:.status,conclusion:.conclusion}]}' || true

Length of output: 110230


@seonghobae Review complete for exact head ea92c326e2dc4e3daa869aff1266c10b05453e7d against protected base 87c4daa1830bac5a5228b6036752ad5633232085. I found no issue in the requested scope.

  • The protected base is an ancestor of the reviewed head.
  • The requested-head commit changes only crates/originweave-core/src/controlled_benchmark.rs.
  • The commit adds 18 rustdoc lines. It does not change executable statements, function signatures, or benchmark decision logic.
  • Every declared production function and method in crates/originweave-core/src/controlled_benchmark.rs now has decision-relevant rustdoc. This includes public API methods, private admission helpers, all three fmt implementations, and ControlledBenchmarkSuiteError::source.
  • The new documentation describes stable run-context comparison order, error formatting, causal error-source exposure, validated-case classification, run-context rejection, and aggregate-counter validation.
  • crates/originweave-core/src/root.rs uses #![deny(missing_docs)]. The added documentation is consistent with that crate contract.
  • The semantic boundaries remain unchanged. The module still derives controlled-suite outcomes from raw trials, rejects malformed evidence, and does not add signing, persistence, browser execution, or broader release authority.

Hosted checks are not green. The requested head has queued checks, and completed checks report skipped. I did not treat either state as passing evidence.

I performed static repository inspection only. I did not run build or test commands.

You are interacting with an AI system.

@seonghobae
seonghobae marked this pull request as ready for review September 14, 2026 22:01
@seonghobae
seonghobae marked this pull request as draft September 14, 2026 23:08

Copy link
Copy Markdown
Contributor Author

Non-conflicting successor security slice is now isolated in Draft #322 instead of moving this exact head. #322 is stacked directly on ea92c326e2dc4e3daa869aff1266c10b05453e7d and adds a test-first fail-closed rule for the 12 Unicode Bidi_Control format characters in product-owned benchmark run-context identities while preserving ordinary visible RTL scripts. This keeps #237's sole current-head queued workflows and completed exact-scope rustdoc review intact. No #237 source/head movement, rerun, cancellation, or gate change was performed.

Copy link
Copy Markdown
Contributor Author

Current execution/repair authority for exact 4d175467c550c969d1ad22e51016473c0a4da034.

The predecessor ea92c326e2dc4e3daa869aff1266c10b05453e7d finally received hosted runners on targeted CI 35655234587 attempt 2 and exposed two real source/test findings rather than an admission problem:

  • Rust contracts 106591222760 failed cargo fmt --all --check in crates/originweave-core/tests/controlled_benchmark_run_context.rs; the hostile-context array needed canonical rustfmt expansion. Python repository contracts passed 152 tests before the formatting failure.
  • Production coverage 106591222543 measured 4,693/4,695 production lines and 5,647/5,650 regions. The only uncovered production lines were crates/originweave-core/src/controlled_benchmark.rs:492-493, the Display arm for ControlledBenchmarkSuiteError::ControlCharacterRunContext. The behavior was exercised, but the error-message contract was not.

Ordinary-forward 4d175467c550c969d1ad22e51016473c0a4da034 repairs both findings in the existing run-context regression: rustfmt-compatible array layout plus a deterministic ControlCharacterRunContext to_string() assertion. No production behavior, threshold, workflow, ruleset, dependency, or release gate was weakened. Compare from ea92c326...: 1 commit / 1 test file / +13/-4 / 0 behind.

New exact-head generation is live: MV3 35689677725, CI 35689677783, Security 35689677811, SAST 35689677835, CodeQL 35689677975, currently queued. Predecessor MV3 GREEN and predecessor security evidence do not transfer to this new head for acceptance.

@coderabbitai review

Please review exact 4d175467c550c969d1ad22e51016473c0a4da034, focusing on whether this test-only repair causally closes the rustfmt and exact owned-production coverage findings without changing benchmark semantics.

@coderabbitai

coderabbitai Bot commented Sep 22, 2026

Copy link
Copy Markdown

Tip

For best results, initiate chat on the files or code changes.

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

base='ea92c326e2dc4e3daa869aff1266c10b05453e7d'
head='4d175467c550c969d1ad22e51016473c0a4da034'
repo='ContextualWisdomLab/OriginWeave'

echo '== exact commits and ancestry =='
git cat-file -t "$base"
git cat-file -t "$head"
git merge-base --is-ancestor "$base" "$head" && echo 'base_is_ancestor=yes'
git rev-list --count "$base..$head"
git diff --check "$base...$head"

echo '== effective changed paths and stat =='
git diff --name-status "$base...$head"
git diff --stat "$base...$head"

echo '== exact repair commit =='
git show --format=fuller --find-renames --find-copies --stat "$head"
git diff --find-renames --find-copies --unified=100 "$base...$head" -- \
  crates/originweave-core/tests/controlled_benchmark_run_context.rs

echo '== test structure =='
ast-grep outline crates/originweave-core/tests/controlled_benchmark_run_context.rs --items all --type function
rg -n -C 12 'ControlCharacterRunContext|to_string\(\)|hostile|control' \
  crates/originweave-core/tests/controlled_benchmark_run_context.rs

echo '== production Display arm and adjacent error contract =='
rg -n -C 20 'ControlCharacterRunContext' \
  crates/originweave-core/src/controlled_benchmark.rs

echo '== production changes in comparison =='
git diff --unified=3 "$base...$head" -- crates/originweave-core/src/controlled_benchmark.rs

echo '== exact-head checks =='
gh api "repos/$repo/commits/$head/check-runs?per_page=100" \
  --jq '.check_runs[] | [.name,.status,(.conclusion // "null")] | `@tsv`' || true
gh pr view 237 --repo "$repo" --json headRefOid,baseRefOid,isDraft,statusCheckRollup \
  --jq '{headRefOid,baseRefOid,isDraft,statusCheckRollup:[.statusCheckRollup[]|{name:.name,status:.status,conclusion:.conclusion}]}' || true

Length of output: 26333


@seonghobae Review complete for exact commit 4d175467c550c969d1ad22e51016473c0a4da034. I found no issue in the requested scope.

  • ea92c326e2dc4e3daa869aff1266c10b05453e7d is an ancestor of the reviewed commit. The comparison contains one commit.
  • The effective diff changes only crates/originweave-core/tests/controlled_benchmark_run_context.rs.
  • The diff contains 13 additions and 4 deletions. It contains no production-source, workflow, dependency, or release-gate change.
  • The hostile-context array now uses rustfmt-compatible multiline layout. This directly addresses the prior formatting finding. I performed static inspection only and did not run cargo fmt --check.
  • control_character_in_reproducibility_context_fails_closed now creates ControlledBenchmarkSuiteError::ControlCharacterRunContext and asserts its full to_string() result.
  • The test performs this assertion for LF, NUL, and U+0085 input. Each input still asserts fail-closed rejection.
  • The asserted string exactly matches the fmt::Display arm at Lines 492-495 in crates/originweave-core/src/controlled_benchmark.rs. This test therefore exercises and verifies the previously uncovered owned-production error-formatting path.
  • The comparison has no change to evaluate_controlled_benchmark_suite_for_run, validate_run_context_field, evidence aggregation, thresholds, or suite outcome logic. The repair does not change benchmark semantics.

The exact-head checks are not passing evidence yet. Rust contracts, Production coverage, MV3, security, SAST, and CodeQL checks are queued. The completed skipped checks do not establish acceptance.

You are interacting with an AI system.

@cwl-noema-review cwl-noema-review Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Noema LLM review

The PR implements a deterministic, fail-closed acceptance policy for the controlled benchmark suite. The implementation strictly enforces the requirement of exactly 100 clean trials for a passing case result and ensures that the overall suite only passes if all required cases (defined by the support profile) pass. Reproducibility contexts are validated byte-for-byte, and trial aggregation prevents evidence fabrication by enforcing unique, one-based ordinals. Documentation and the comprehensive test suite accurately reflect these constraints.

Reviewed changed lines

  • crates/originweave-core/src/controlled_benchmark.rs:564 (RIGHT): The implementation of evaluate_controlled_benchmark_case correctly enforces the deterministic fail-closed logic. Probe 1: line 565 explicitly returns Err(ControlledBenchmarkError::NonCanonicalTrialCount) if total_trials exceeds CONTROLLED_DETERMINISTIC_REQUIRED_TRIALS. Probe 2: lines 605-611 verify that if total_trials < 100 and no failures are present, it returns ControlledBenchmarkCaseOutcome::Inconclusive.
  • crates/originweave-core/src/controlled_benchmark.rs:665 (RIGHT): The suite evaluation logic correctly handles the support profile and registry. Probe 1: lines 666-668 return Err(ControlledBenchmarkSuiteError::InvalidSupportProfile) if native_messaging is true while manifest_v3 is false. Probe 2: lines 718-721 identify missing required cases via missing_required check, which subsequently triggers BenchmarkSuiteOutcome::Inconclusive at line 725.
  • crates/originweave-core/src/controlled_benchmark.rs:635 (RIGHT): The reproducibility context validation is strictly fail-closed. Probe 1: line 755 (trimmed != value) correctly triggers NonCanonicalRunContext to reject surrounding whitespace. Probe 2: line 758 (char::is_control) correctly triggers ControlCharacterRunContext for any control characters.
  • crates/originweave-core/src/controlled_benchmark.rs:416 (RIGHT): The aggregate_controlled_benchmark_trials function prevents evidence fabrication. Probe 1: lines 426-430 use a BTreeSet to ensure DuplicateTrialOrdinal is returned if an ordinal is repeated. Probe 2: lines 421-425 explicitly check if trial_ordinal is within 1..=100, returning InvalidTrialOrdinal otherwise.
  • ARCHITECTURE.md:80 (RIGHT): Documentation accurately reflects implementation. ARCHITECTURE.md (lines 80-83) now describes the versioned benchmark case identities and threshold evaluation. CHANGELOG.md (line 14) explicitly documents the 100-clean-trial requirement and the fail-closed behavior for missing/inconclusive cases.

Adversarial validation

  • crates/originweave-core/src/controlled_benchmark.rs:565 (RIGHT) falsified: Providing more than 100 trials might allow a case to pass despite failures via dilution. — The logic at line 565 explicitly returns Err(ControlledBenchmarkError::NonCanonicalTrialCount) if total_trials > 100, rejecting the evidence entirely.
  • crates/originweave-core/src/controlled_benchmark.rs:666 (RIGHT) falsified: A support profile could claim native messaging without Manifest V3, leading to an invalid state. — Lines 666-668 explicitly check this condition and return Err(ControlledBenchmarkSuiteError::InvalidSupportProfile).
  • Residual risk: None. The logic is purely deterministic and operates on raw evidence, making it resistant to typical runtime behavioral regressions as long as the evidence pipeline provides correct raw trial data.

Findings

  • No blocking findings.
  • Result: APPROVE
  • Head SHA: 4d175467c550c969d1ad22e51016473c0a4da034
  • Reviewer credential: noema-review-github-app-refresh
  • Actor: cwl-noema-review[bot]

@opencode-agent opencode-agent Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

OpenCode reviewed the current-head product diff. Coverage is a separate gate.

Changed files

  • ARCHITECTURE.md — repository behavior
  • CHANGELOG.md — repository behavior
  • crates/originweave-core/src/controlled_benchmark.rs — Rust workspace crate API and tests
  • crates/originweave-core/src/root.rs — Rust workspace crate API and tests
  • crates/originweave-core/tests/controlled_benchmark.rs — Rust workspace crate API and tests
  • crates/originweave-core/tests/controlled_benchmark_registry_version.rs — Rust workspace crate API and tests
  • crates/originweave-core/tests/controlled_benchmark_run_context.rs — Rust workspace crate API and tests
  • crates/originweave-core/tests/controlled_benchmark_suite_authority.rs — Rust workspace crate API and tests
  • crates/originweave-core/tests/controlled_benchmark_support_profile.rs — Rust workspace crate API and tests
  • crates/originweave-core/tests/controlled_benchmark_trial_aggregation.rs — Rust workspace crate API and tests
  • crates/originweave-core/tests/controlled_benchmark_unauthorized_side_effect_count.rs — Rust workspace crate API and tests
  • crates/originweave-core/tests/controlled_benchmark_zero_trial_side_effect.rs — Rust workspace crate API and tests

Changed behavior

classDiagram
  class ControlledBenchmarkCaseId
  class ControlledBenchmarkSupportProfile
  class ControlledBenchmarkRunContext
  class ControlledBenchmarkCaseOutcome
  class ControlledBenchmarkCaseEvidence
  class ControlledBenchmarkTrialEvidence
  class ControlledBenchmarkCaseTrials
  class ControlledBenchmarkTrialAggregationError
Loading

Changed API

  • ControlledBenchmarkCaseId
  • ControlledBenchmarkSupportProfile
  • ControlledBenchmarkRunContext
  • ControlledBenchmarkCaseOutcome
  • ControlledBenchmarkCaseEvidence
  • ControlledBenchmarkTrialEvidence
  • ControlledBenchmarkCaseTrials
  • ControlledBenchmarkTrialAggregationError
  • aggregate_controlled_benchmark_trials
  • ControlledBenchmarkError
  • ControlledBenchmarkSuiteError
  • evaluate_controlled_benchmark_case
  • evaluate_controlled_benchmark_suite_for_run
  • evaluate_controlled_benchmark_suite
  • controlled_benchmark
  • mcp
  • release_acceptance

Findings

No source-backed product finding is synthesized from the coverage gate. A coverage miss belongs in the status comment.

  • Head SHA: 4d175467c550c969d1ad22e51016473c0a4da034
  • Workflow run: 35714835267
  • Workflow attempt: 1
  • Coverage gate: failure

Review outcome

Coverage is a gate, not the review. This body reviews the changed product files.

Changed-File Evidence Map

classDiagram
  class ControlledBenchmarkCaseId
  class ControlledBenchmarkSupportProfile
  class ControlledBenchmarkRunContext
  class ControlledBenchmarkCaseOutcome
  class ControlledBenchmarkCaseEvidence
  class ControlledBenchmarkTrialEvidence
  class ControlledBenchmarkCaseTrials
  class ControlledBenchmarkTrialAggregationError
Loading

@opencode-agent

opencode-agent Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

OpenCode Review Overview

Coverage evidence did not pass, so approval is blocked. The formal pull-request review is the source-backed diff review, not this status comment.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant