Repository navigation
decider: reuse request prefixes with conservative auto selection - #113
Levius-Fubuki wants to merge 24 commits into
Conversation
Add CUDA operations that continue a sequence after a shared prefix, for request-local prefix reuse (ThinkFlowLab#85): the Gated DeltaNet conv with the inputs of the three positions before it, the chunked prefill from and to a float32 state, gated attention for queries after cached keys, and a pitched device-to-device row copy. The existing conv, prefill and gated attention are these operations without history, state or cached positions. Bump the ABI to 6 and add the Rust bindings and GPU tests against the unsplit calls.
cuBLASLt's heuristic picks an algorithm per M, so a row's result can change with the number of rows in the call, and a prefix and its branch run separately round differently from one pass over the same tokens. Add cs1_gemm_create_fixed: one algorithm per weight shape (N, K, ldy), the heuristic's first choice at a reference M among algorithms without split-K, for every M. The default handle is unchanged. Bump the ABI to 7 and add a GPU test that each projection's rows match across M.
hsliuustc0106
left a comment
There was a problem hiding this comment.
Reviewed the cumulative tip at head d31bfd4a5e (24 commits; merge base 4a79980d), including the new src/models/decider/native crate, the shared qwen3_5 executor changes, and the ABI-7 CUDA kernels. No P0/P1/P2 findings — details below; two P3s and readiness gaps that block merge per CONTRIBUTING.
Verified in code: ABI lockstep holds (ops.h CS1_ABI_VERSION 7 = cuda.rs ABI_VERSION: u32 = 7; every new C symbol — cs1_copy_rows, cs1_gdn_conv_history, cs1_gdn_prefill_state, cs1_attention_gated_cached, cs1_gemm_create_fixed — has a matching Rust api! declaration; a stale library is rejected at load with a rebuild hint). The generalized flash-attention kernel keeps per-query accumulation bit-identical to the unsplit call (key tiles walk in the same order; masked entries don't disturb the running max); cs1_gdn_conv_history with null history is bit-identical to the legacy kernel; the stateful GDN prefill's float2 accesses are aligned and the 64-token split preserves unsplit chunk boundaries; cs1_gemm_create_fixed hard-errors rather than silently switching when an M can't be served, and the header honestly discloses that M-independence is device-checked (tests/qwen3_5/kernels.rs). The request-local prefix plan preserves row order, question identity, and token accounting: prefixes end on GDN_CHUNK=64 boundaries, unaligned remainders fold into group prefixes, branches run serially on one stream over snapshotted state and write only their own KV rows, common-prefix computation excludes each row's final token (non-empty branches), and bit-exact-vs-forward_fixed parity is asserted in opt-in GPU tests. The conservative auto selector matches the body (saved ≥ 4096 && saved ≥ ⌈total/3⌉), and DECIDER_PREFIX/DECIDER_FIXED/DECIDER_GRAPH mutual exclusion is enforced in three places. On execution error both streams synchronize and the model retires before permit release.
P3 findings:
tests/decider/data/prefix-workloads.json— this retained corpus has no in-tree consumer reference: thetests/decider/data/README.mddocumenting the data directory was deleted in this same series, so the corpus's role is discoverable only from out-of-repo archives. One line in a retained doc (or a comment inverify_reference.py) naming it would fix this.- The "Full self-review completed" sections duplicated across the #110–#113 bodies are pinned to superseded snapshot heads with stale counts. Please correct or remove them.
Readiness gaps (required before merge): this PR crosses the large-change threshold (5,236 authored / 5,516 total at head, per my count) but its description has no self-review checklist at all (the section reads only "Maintainer review remains outstanding."), no component map or suggested review order (only the merge-order list; the per-component narrative lives in the #110–#112 bodies), doesn't state the total diff size separately from the authored count, and its validation summary is a single paragraph rather than per affected area (ABI/kernel boundary, shared-executor changes, decider integration).
Documentation gap: the series deleted all decider documentation that existed at intermediate heads (model contract README, recipe/decider/README.md and validation docs, tests/decider/data/README.md, and the docs/supported-models.md/mkdocs.yml/root-README entries). CONTRIBUTING requires these for a new model; merging as-is would leave the model undocumented on the docs site and unsupported-models list, and also causes the corpus P3 above.
Deployment note: ABI 7 means libqwen3_5_cuda.so and all consumers (Cua-S1, Open-Jev included) must be rebuilt and restarted together; a stale deployment fails at load, not silently.
CI (rust/build/benchmarks) passes on this and all four series heads (observed). All GPU and pinned-reference parity results remain author-reported at snapshot heads (archives not independently hash-checked or executed); no GPU was used and no Rust toolchain exists on this host.
Provenance: canonical .agents/skills/system1-omni-review/SKILL.md (SHA-256 58be3bc6…dde108) and its repository map read at trusted base 4a79980d8a75, plus CONTRIBUTING.md, docs/architecture.md, and src/runtime/README.md. Remote head rechecked before posting.
Purpose
Repeated Decider rows can share long request/question prefixes, while forcing short rows through fixed-GEMM execution roughly doubles their latency. Add request-local prefix reuse with an opt-in conservative selector:
DECIDER_PREFIX=autouses shared execution only when the aligned plan saves at least 4096 tokens and at least one third of original row-token work; otherwise it keeps normal independent/packed eager execution.DECIDER_PREFIX=1andDECIDER_FIXED=1retain forced behavior. Prefix/fixed/auto reject Graph combinations. Request-local branch state stays independent and responses/calibration/usage are preserved. Merge order: #94 → #110 → #111 → #112 → #113, with shared-Qwen #97/#98/#99 prerequisites before prefix execution. Upstream prerequisite commits retain their authors. ABI 7 requires rebuilding the backend and consumers together.Submission scope
Core implementation and directly related tests are retained. Ordinary documentation, experiment reports/media and generated run outputs remain outside this submission. Test fixtures, build inputs, deployment data and third-party licensing remain as needed. The diagnostic test helper lives at
tests/decider/runner.rs, built withcargo build -p omni-decider-native --example decider-run --release --locked; its executable istarget/release/examples/decider-run. It is not a production worker binary.Validation
After cleaning workspace debug artifacts before each branch, fresh local fmt, Clippy/all-targets with warnings denied and 117 Rust tests pass (31 opt-in tests ignored). Independent GitHub CI passes on this exact head, including release workspace build and strict Docs. Restored test bodies/fixtures match the saved originals. Production execution is unchanged. GPU tests were not rerun; the server remains shut down. 5 Python regression tests also pass.
Head
d31bfd4a5e7a3cbecc9653c551a54b110d87fb94; 56 changed files and 5,236 changed source/test/build lines, excluding licenses, notices, locks and static data.Self-review
Maintainer review remains outstanding.