Skip to content

feat(decider): bounded request-local CUDA batching - #111

Open
Levius-Fubuki wants to merge 10 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/decider-request-batching
Open

Levius-Fubuki wants to merge 10 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/decider-request-batching

Conversation

@Levius-Fubuki

@Levius-Fubuki Levius-Fubuki commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Pack contiguous independent Decider rows within one admitted request and project their selected BF16 label logits in one GEMM. Defaults preserve one-row eager execution. DECIDER_BATCH_MAX_ROWS=1..4 and DECIDER_BATCH_MAX_TOKENS=1..4096 bound each pack; longer rows run alone. Response reconstruction, Score normalization, usage, and whole-request synchronization/retirement remain unchanged.

Response reconstruction, Score normalization, usage and whole-request synchronization/retirement are preserved. Depends on #110 and #94; the main-based diff includes their core implementations.

Submission scope

Core implementation and directly related tests are retained. Ordinary documentation, experiment reports/media and generated run outputs remain outside this submission. Test fixtures, build inputs, deployment data and third-party licensing remain as needed. The diagnostic test helper lives at tests/decider/runner.rs, built with cargo build -p omni-decider-native --example decider-run --release --locked; its executable is target/release/examples/decider-run. It is not a production worker binary.

Validation

After cleaning workspace debug artifacts before each branch, fresh local fmt, Clippy/all-targets with warnings denied and 108 Rust tests pass (20 opt-in tests ignored). Independent GitHub CI passes on this exact head, including release workspace build and strict Docs. Restored test bodies/fixtures match the saved originals. Production execution is unchanged. GPU tests were not rerun; the server remains shut down. 2 Python regression tests also pass.

Head 7eea74a9776fdce97917c6c9eab14da67c19c596; 35 changed files and 2,851 changed source/test/build lines, excluding licenses, notices, locks and static data.

Self-review

Full focused diff reviewed locally and by an independent reviewer. Generated logs and measurements remain outside the committed source tree. Requested full agent-assisted self-review is complete; maintainer review remains separate. See CONTRIBUTING.md for the full checklist.

Main-based diff:2847authored lines including manifests,3529total; focused batching commit19b262e5.

Final followup evidence

Immutable sequential validation archive, asset decider-next-20261007-final-evidence.tar.gz; SHA256 9f57d969fc484c518d8ebac7b2dd38174e2c6954e7af04b9352550675053ae48. Includes exact5branch heads, final source-vs-executed-runtime audit, frozen protocols, raw repeated samples/responses, failures, source/binary/library hashes, environment and full self-review. Weights/credentials/private inputs excluded.

Final heads: contract f59cca40462d5b63c3dbc909e6aefb147bd8f44e; worker 9a09e5a58d33692edd5aab5634f268055e4c8151; batching 19b262e544de2c9e2d0b23f865204a6ae2372038; Graph d88179c6f3b78cd33025031134b7a8086e61f481; dependent prefix 6962971820002d46b030e09503cd21d44c831083. Main base 99865743d27316fbe81362dc0f0e6de6fda86284. All current-head CI Rust/benchmark/docs checks passed; docs deploy intentionally skipped on PRs. Runtime source/manifests match measured snapshots byte-for-byte; final prefix head adds a documentation-only command repair.

2026-10-08 verification follow-up

Validated source head: 02efb1b4a6e7c546aacefc140c0632461dab48db. Upstream main snapshot: 4a79980d; GitHub reports no merge conflict. Fresh exact-head fmt, strict workspace Clippy, workspace tests and release build pass: 108 passed, 0 failed, 20 ignored. Ignored tests are counted separately from explicit GPU execution. Exact commands and logs are in the archive. Merged current main Open-Jev-9B support and resolved documentation/index conflicts while retaining both model entries. Strict MkDocs also passes. GPU/reference results above were executed on snapshot 19b262e544de2c9e2d0b23f865204a6ae2372038; a complete source comparison confirms all affected model/runtime/backend and regression sources are byte-identical at this new integration head (see gpu-source-binding.json). Open-Jev-9B inference is not newly validated by these GPU cases. Fresh GPU execution passes the 3 complete-request/batching regressions and 2 BF16 head tests. The eager-only branch uses CUA_S1_GRAPH=0; an initial helper used 1 and was correctly rejected before inference, with the failure preserved. Supported packing remains capped at 4. Keep dependent on #110/#94.

Raw verification evidence, SHA-256 5df66bf3974d534875946168fba4d7fe543d08edbd658c3756f17d6f0c8f5773. Includes exact source manifests, environment/library/binary hashes, complete logs, synthetic reference/HTTP results and initial setup failures. CUDA execution is serial on one RTX 4090. No weights, credentials or private inputs are included. Contributor/maintainer review and dependency merges remain separate.

Completed self-review checklist

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

Full self-review completed / Ready — 2026-10-08

Current head: 354985918cf93cd84d706d85cba2deb1db2dff04; merge base: 4a79980d8a75fd063cb3f8247e06215e288b18ac. Full cumulative diff: 2847 authored-code lines / 3529 total diff lines. The author explicitly requested completed self-review and Ready conversion. The completed checklist records this agent-assisted review and its verification; it does not represent maintainer approval.

Root integration review plus three independent component reviews covered all changed source, tests, helpers, manifests and documentation: CPU contract/checkpoint/tokenizer/calibration; serial worker/admission/GPU head and error retirement; ordered bounded packing and response reconstruction; Graph lifetime/cache/health/options; ABI7 CUDA continuations/fixed GEMM/request-local prefix state; root regressions, recipe reproduction, fixture consumers and evidence/claim boundaries. No functional blocker remains. For #112/#113 above 3,000 authored lines, this is full component/integration self-review, with the existing split rationale and review order retained.

Corrected stale unsupported-Graph and no-worker-prefix-consumer documentation where applicable. Removed unused CPU fixture run/duplicate fields on the Decider stack, preserving all 255 label IDs and all 12 used fixture values exactly; the original removed data remain in evidence outside committed source. Both pinned CPU contract opt-ins pass on every repaired Decider branch; repaired docs pass strict MkDocs. All five initial review heads passed fresh fmt, strict workspace Clippy, workspace tests, release build and strict docs; latest-head CI Rust/benchmark/docs also passes. Ignored hardware/model tests are separate from executed tests.

Runtime/CUDA/GPU regression code is unchanged by this cleanup. Earlier personally generated GPU/reference/HTTP campaigns retain their original source revisions; independent reviewers rehashed both source bindings, replayed all stored numerical gates and fixed/shared exact controls, and recomputed published timings. ready-source-binding.json binds these runtime checks to the latest heads and separately proves the used CPU oracle values are unchanged. No new CUDA/HTTP/performance campaign was run after server shutdown; the server remains off. Existing hardware, scope and performance limitations remain disclosed.

Complete self-review reports and fresh checks; SHA-256 1edbd4a6e13576d4cc23638b7ec1e35c1248d6a712b8a6406e1f269ca685cb9d. Includes all four review reports, closure summary, original reviewed full diffs, new CPU/docs logs, latest-head CI, source/fixture binding and archived removed fields. Merge dependencies remain #103 before #106, and #94 → #110 → #111 → #112 → #113, including #97/#98/#99 prerequisites for #113; Ready conversion does not merge them.

@Levius-Fubuki
Levius-Fubuki marked this pull request as ready for review October 8, 2026 01:11
Copilot AI balanced review requested due to automatic review settings October 8, 2026 01:11

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at head 7eea74a977 (10 commits; step delta over #110's head). No actionable findings.

The batching delta is sound: BatchLimits clamps to rows 1..=4 and tokens 1..=4096; ranges() does greedy ordered packing where an over-budget row runs alone (verified for over-budget and overflow edges with checked_add/saturating_add); Head::project_batch uses capacity-sized device buffers with per-row finite/count validation and per-row BF16 output parse; forward_batch packs multi-row ranges inside the single admitted request unit. Row order and per-row counts are unit-tested, defaults (1 row) preserve #110 behavior, and the env parsing matches the body. Request order preservation and whole-request admission semantics hold.

Template, checklist, and stated counts (2,851 authored / 3,077 total — both match my count) are in order. One P3: the body's "Full self-review completed" section still pins superseded head 354985918c with 2847/3529 counts (the Validation section is current) — please update or drop the stale section.

CI (rust/build/benchmarks) passes on this head (observed). The author-reported GPU batch-parity results at snapshot heads were not independently executed by this review — no GPU was used and no Rust toolchain exists on this host.

Provenance: canonical .agents/skills/system1-omni-review/SKILL.md (SHA-256 58be3bc6…dde108) and its repository map read at trusted base 4a79980d8a75, plus CONTRIBUTING.md and docs/architecture.md. Remote head rechecked before posting.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants