Skip to content

decider: reuse request prefixes with conservative auto selection - #113

Open
Levius-Fubuki wants to merge 24 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/decider-shared-prefix
Open

Levius-Fubuki wants to merge 24 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/decider-shared-prefix

Conversation

@Levius-Fubuki

@Levius-Fubuki Levius-Fubuki commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Repeated Decider rows can share long request/question prefixes, while forcing short rows through fixed-GEMM execution roughly doubles their latency. Add request-local prefix reuse with an opt-in conservative selector: DECIDER_PREFIX=auto uses shared execution only when the aligned plan saves at least 4096 tokens and at least one third of original row-token work; otherwise it keeps normal independent/packed eager execution.

DECIDER_PREFIX=1 and DECIDER_FIXED=1 retain forced behavior. Prefix/fixed/auto reject Graph combinations. Request-local branch state stays independent and responses/calibration/usage are preserved. Merge order: #94 → #110 → #111 → #112 → #113, with shared-Qwen #97/#98/#99 prerequisites before prefix execution. Upstream prerequisite commits retain their authors. ABI 7 requires rebuilding the backend and consumers together.

Submission scope

Core implementation and directly related tests are retained. Ordinary documentation, experiment reports/media and generated run outputs remain outside this submission. Test fixtures, build inputs, deployment data and third-party licensing remain as needed. The diagnostic test helper lives at tests/decider/runner.rs, built with cargo build -p omni-decider-native --example decider-run --release --locked; its executable is target/release/examples/decider-run. It is not a production worker binary.

Validation

After cleaning workspace debug artifacts before each branch, fresh local fmt, Clippy/all-targets with warnings denied and 117 Rust tests pass (31 opt-in tests ignored). Independent GitHub CI passes on this exact head, including release workspace build and strict Docs. Restored test bodies/fixtures match the saved originals. Production execution is unchanged. GPU tests were not rerun; the server remains shut down. 5 Python regression tests also pass.

Head d31bfd4a5e7a3cbecc9653c551a54b110d87fb94; 56 changed files and 5,236 changed source/test/build lines, excluding licenses, notices, locks and static data.

Self-review

Maintainer review remains outstanding.

Levius-Fubuki and others added 14 commits October 7, 2026 20:50
Add CUDA operations that continue a sequence after a shared prefix,
for request-local prefix reuse (ThinkFlowLab#85): the Gated DeltaNet conv with the
inputs of the three positions before it, the chunked prefill from and
to a float32 state, gated attention for queries after cached keys, and
a pitched device-to-device row copy. The existing conv, prefill and
gated attention are these operations without history, state or cached
positions. Bump the ABI to 6 and add the Rust bindings and GPU tests
against the unsplit calls.
cuBLASLt's heuristic picks an algorithm per M, so a row's result can
change with the number of rows in the call, and a prefix and its branch
run separately round differently from one pass over the same tokens.
Add cs1_gemm_create_fixed: one algorithm per weight shape (N, K, ldy),
the heuristic's first choice at a reference M among algorithms without
split-K, for every M. The default handle is unchanged. Bump the ABI to 7
and add a GPU test that each projection's rows match across M.
@twu3202

twu3202 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Rebased #97 to #99 onto main (after #102 and #92). #99 now splits the layer loop the way your merge does, packed sequences plus single-sequence spans, so #113 can build on its head directly.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@Levius-Fubuki

Copy link
Copy Markdown
Collaborator Author

Checked the rebased #99: the packed/span loop is compatible, and the CUDA kernels/FFI match #113. I’ll sync #113 onto the new head while preserving its Graph stats, explicit mode and lifetime handling.

@Levius-Fubuki Levius-Fubuki changed the title feat(decider): request-local shared prefix reuse decider: reuse request prefixes with conservative auto selection Oct 8, 2026

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the cumulative tip at head d31bfd4a5e (24 commits; merge base 4a79980d), including the new src/models/decider/native crate, the shared qwen3_5 executor changes, and the ABI-7 CUDA kernels. No P0/P1/P2 findings — details below; two P3s and readiness gaps that block merge per CONTRIBUTING.

Verified in code: ABI lockstep holds (ops.h CS1_ABI_VERSION 7 = cuda.rs ABI_VERSION: u32 = 7; every new C symbol — cs1_copy_rows, cs1_gdn_conv_history, cs1_gdn_prefill_state, cs1_attention_gated_cached, cs1_gemm_create_fixed — has a matching Rust api! declaration; a stale library is rejected at load with a rebuild hint). The generalized flash-attention kernel keeps per-query accumulation bit-identical to the unsplit call (key tiles walk in the same order; masked entries don't disturb the running max); cs1_gdn_conv_history with null history is bit-identical to the legacy kernel; the stateful GDN prefill's float2 accesses are aligned and the 64-token split preserves unsplit chunk boundaries; cs1_gemm_create_fixed hard-errors rather than silently switching when an M can't be served, and the header honestly discloses that M-independence is device-checked (tests/qwen3_5/kernels.rs). The request-local prefix plan preserves row order, question identity, and token accounting: prefixes end on GDN_CHUNK=64 boundaries, unaligned remainders fold into group prefixes, branches run serially on one stream over snapshotted state and write only their own KV rows, common-prefix computation excludes each row's final token (non-empty branches), and bit-exact-vs-forward_fixed parity is asserted in opt-in GPU tests. The conservative auto selector matches the body (saved ≥ 4096 && saved ≥ ⌈total/3⌉), and DECIDER_PREFIX/DECIDER_FIXED/DECIDER_GRAPH mutual exclusion is enforced in three places. On execution error both streams synchronize and the model retires before permit release.

P3 findings:

  • tests/decider/data/prefix-workloads.json — this retained corpus has no in-tree consumer reference: the tests/decider/data/README.md documenting the data directory was deleted in this same series, so the corpus's role is discoverable only from out-of-repo archives. One line in a retained doc (or a comment in verify_reference.py) naming it would fix this.
  • The "Full self-review completed" sections duplicated across the #110–#113 bodies are pinned to superseded snapshot heads with stale counts. Please correct or remove them.

Readiness gaps (required before merge): this PR crosses the large-change threshold (5,236 authored / 5,516 total at head, per my count) but its description has no self-review checklist at all (the section reads only "Maintainer review remains outstanding."), no component map or suggested review order (only the merge-order list; the per-component narrative lives in the #110–#112 bodies), doesn't state the total diff size separately from the authored count, and its validation summary is a single paragraph rather than per affected area (ABI/kernel boundary, shared-executor changes, decider integration).

Documentation gap: the series deleted all decider documentation that existed at intermediate heads (model contract README, recipe/decider/README.md and validation docs, tests/decider/data/README.md, and the docs/supported-models.md/mkdocs.yml/root-README entries). CONTRIBUTING requires these for a new model; merging as-is would leave the model undocumented on the docs site and unsupported-models list, and also causes the corpus P3 above.

Deployment note: ABI 7 means libqwen3_5_cuda.so and all consumers (Cua-S1, Open-Jev included) must be rebuilt and restarted together; a stale deployment fails at load, not silently.

CI (rust/build/benchmarks) passes on this and all four series heads (observed). All GPU and pinned-reference parity results remain author-reported at snapshot heads (archives not independently hash-checked or executed); no GPU was used and no Rust toolchain exists on this host.

Provenance: canonical .agents/skills/system1-omni-review/SKILL.md (SHA-256 58be3bc6…dde108) and its repository map read at trusted base 4a79980d8a75, plus CONTRIBUTING.md, docs/architecture.md, and src/runtime/README.md. Remote head rechecked before posting.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants