diff --git a/.agents/skills/self-review/SKILL.md b/.agents/skills/self-review/SKILL.md index e6627d1..59bf4dd 100644 --- a/.agents/skills/self-review/SKILL.md +++ b/.agents/skills/self-review/SKILL.md @@ -10,6 +10,15 @@ instructions. Record the actual target branch and base/head commits, and review the full diff from their merge base plus relevant uncommitted changes. Distinguish what is in the PR from local-only work; disclose if the base could not be refreshed. +Apply the [large-code-change requirements](../../../CONTRIBUTING.md#large-code-changes): +report authored-code and total diff counts separately using the guide's counting +convention. Above 3,000 changed code lines, verify the contributor's full +self-review, split rationale, component map/review order, and validation across +affected components and interfaces before recommending readiness. A quick +precheck is insufficient; keep the PR draft until the contributor self-review is +complete. Report missing preparation as a readiness gap; size alone is not a +correctness finding or a reason to require expensive model/GPU runs. + Check correctness, focused scope, decision-model and front contracts, fallback behavior, cancellation/timeouts, and resource cleanup as relevant. Verify tests cover the changed behavior, including failure paths and a regression case for a @@ -23,6 +32,28 @@ This skill prepares a local contributor report. It does not itself authorize edits, commits, pushes, external posts, paid model calls, downloads, or changes to review status. Use only separately authorized execution resources and budgets. +## Committed artifact hygiene + +Apply this check in every review, including quick prechecks. Inspect added and +changed artifacts in the complete diff, including JSON/JSONL, CSV, logs, reports, +source/binary hash inventories, and generated media. Classify them by purpose and +actual consumer, rather than rejecting a file extension. + +- Keep necessary configuration, request examples, maintained benchmark inputs, + and small deterministic fixtures or reference oracles in the repository's + intended locations. Identify the test, tool, or documented workflow that needs + each retained artifact, such as labelled agent datasets or replay inputs. +- Flag one-off run summaries, response dumps, cache statistics, profiler output, + agent process notes, and duplicate historical results that have no maintained + source-tree role. A link from PR prose or documentation alone does not justify + committing generated run output. Report concrete paths and consumers. +- Preserve raw measurements, failures, and provenance in a durable artifact + archive or PR/CI evidence, and link the exact revision or run from the summary. + Do not discard evidence to reduce the diff or hide it in a committed archive. +- When removing redundant output, check its callers, links, and reproduction + commands. Keep replay inputs and expected responses intact; verify their + hashes and rerun the affected replay or documentation checks. + ## Task evidence For changes to what an agent can accomplish, show a concrete task as **input → diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index d1437b2..5ace7f3 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -60,6 +60,33 @@ week for Python packages and one for the GitHub Actions. says how to verify. Include a **Demo / evidence** section; use the [self-review skill](.agents/skills/self-review/SKILL.md) and the recording guide below. +Apply [committed artifact hygiene](.agents/skills/self-review/SKILL.md#committed-artifact-hygiene) +in every review, including quick prechecks. Keep maintained fixtures and replay inputs; preserve raw run +evidence in durable, reviewer-accessible PR/CI artifacts rather than committing redundant generated output. + +### Large code changes + +PRs with **more than 3,000 changed lines of authored code** need extra contributor +attention before requesting review. Count additions plus deletions against the +PR's merge base in source files, tests, and build or validation scripts. Report +this count separately from the total diff size; exclude documentation, generated +output, lockfiles, and static fixtures from the code count, while still reviewing +those files for relevance and correctness. + +- Complete a full self-review of every affected component and its integration + boundaries. A quick precheck alone is insufficient; keep the PR in draft until + the contributor self-review is complete. +- Consider splitting independent features, refactors, and cleanup into focused + PRs. If the change needs to stay together, explain why in the PR description + and provide a component map and suggested review order. +- Include the code-line count and a validation summary for each affected area in + the PR description: commands, results, and unverified behavior with reasons. + Cover changed interfaces between components as well as individual components. + +Size signals the need for closer review; it is not itself a correctness finding. +Choose checks based on the changed behavior and risk. Crossing this threshold +alone does not require GPU benchmarks or other expensive experiments. + ## Add an agent use-case recipe Use [recipes/README.md](recipes/README.md) and [the template](recipes/TEMPLATE.md) to document a complete task @@ -126,8 +153,10 @@ Export a readable MP4 with H.264 where possible; aim below 10 MB. GitHub's [attachment guide](https://docs.github.com/en/get-started/writing-on-github/working-with-advanced-formatting/attaching-files) lists supported formats and current limits. Drag the clip into the PR description's **Demo / evidence** section or a PR comment, wait for the upload to finish, and save the resulting link. Uploading makes the file public for -this public repository. Use attachments for videos; keep sanitized reproduction commands and small result -records in the repository or a durable reviewer-accessible archive. Verify the uploaded video plays and that +this public repository. Use attachments for videos; keep sanitized reproduction commands and maintained +fixtures in the repository, and generated run records in a durable reviewer-accessible archive or PR/CI +evidence, following [artifact hygiene](.agents/skills/self-review/SKILL.md#committed-artifact-hygiene). +Verify the uploaded video plays and that reviewers can open its linked trace. Use a caption such as: diff --git a/s1a/decision_models/served.py b/s1a/decision_models/served.py index 50a1b4d..d9737c2 100644 --- a/s1a/decision_models/served.py +++ b/s1a/decision_models/served.py @@ -304,6 +304,7 @@ class ServedLayaModel(DecisionModel): name = "laya-served" deterministic = True + bills_input_tokens = False # the client's own server, not Jev's pricing def __init__( self, diff --git a/tests/test_decision_models_base.py b/tests/test_decision_models_base.py index e0dab1c..ff3539d 100644 --- a/tests/test_decision_models_base.py +++ b/tests/test_decision_models_base.py @@ -23,6 +23,7 @@ RuleModel, ScriptedModel, ) +from s1a.decision_models.served import ServedLayaModel class TestBillingContract(TestCase): @@ -30,6 +31,7 @@ def test_every_backend_declares_whether_input_tokens_use_jev_pricing(self) -> No expected = { JevModel: True, LayaModel: False, + ServedLayaModel: False, CuaS1Model: False, RandomModel: False, RuleModel: False, diff --git a/tests/test_tool_models.py b/tests/test_tool_models.py index a8e002a..e73f299 100644 --- a/tests/test_tool_models.py +++ b/tests/test_tool_models.py @@ -198,6 +198,7 @@ class ServedStub(DecisionModel): """Answers the way the served model does: an identity in ``model``, the server's facts in ``raw``.""" name = "laya-served" + bills_input_tokens = False @property def model(self) -> str: