Repository navigation
[Docs] Require artifact hygiene and full review for large code changes - #47
Merged
Merged
Conversation
hsliuustc0106
marked this pull request as ready for review
October 6, 2026 23:45
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Port the review safeguards from merged ThinkFlowLab/system1-omni#105 to System1-Agents: redundant generated output should not obscure a change, and large authored-code diffs need a complete contributor review.
How
CONTRIBUTING.mddefines the >3,000 changed authored-code line threshold: additions + deletions from the merge base in source, tests, and build/validation scripts. Documentation, generated output, lockfiles and static fixtures stay out of that count but remain in review scope.What
Two Markdown files change, plus the isolated three-file CI repair below. Authored-code count: 4 additions + 0 deletions. Total diff: 66 additions + 2 deletions = 68 changed lines. Base/merge base:
621681b5d59a779e3e93b9524c1772a1970fe7d7. Head:945dc4e8e637fbfc5ae6eedbd8f19d55c717ac2e.Isolated CI repair
The original documentation head inherited three failures already present on exact main
621681b:ServedLayaModeland theServedStubtest double omittedbills_input_tokens, causing billing-contract failures and an AttributeError in the tool wrapper. The earlier main CI run and original PR run show the same failures.A separate follow-up commit reuses only the focused correction from #37 at
13d9613d896c9ee8a8f663614e5c334d2ccdab09: explicitly declare both served classes non-billable at Jev rates and add the real served backend to the billing-contract expected map. Existing contract and tool-provenance tests cover the regression. No Laya dependency, precision or other inference changes from #37 are included.PR #45's video-evidence edits are left intact and separate. Its overlapping CONTRIBUTING/self-review patch applies cleanly on top of this change; no video requirement is weakened.
Verification
git diff --checkagainst exact source contents.git apply --check --include=CONTRIBUTING.md --include=.agents/skills/self-review/SKILL.md ../agents-pr45.patchfor the pending [Docs] Require application, agent and Omni video demos #45 overlap.Falseafter it; the served expectation is present in the regression contract. These focused checks are not a full pytest run.945dc4e8e637fbfc5ae6eedbd8f19d55c717ac2e: all three core jobs (Linux 3.11/3.13, Windows 3.11) passed with 606 tests passed and 43 skipped each; full passed with 615 tests passed and 33 skipped. Smoke passed in all four jobs, along with configured lint/type/build checks. Browser is intentionally skipped for pull-request events. No failed checks remain in this run.Demo / evidence
No generated run artifacts or media are added or removed. The policy portion needs no agent demo. The four-line CI repair restores an existing served-model billing contract and is validated by automated regression tests; no live inference/performance claim or model/GPU run is made.