Skip to content

decider: add verified eager native CUDA worker - #110

Open
Levius-Fubuki wants to merge 11 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/decider-native-worker
Open

Levius-Fubuki wants to merge 11 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/decider-native-worker

Conversation

@Levius-Fubuki

@Levius-Fubuki Levius-Fubuki commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Decider-2B v11 currently has CPU request compilation/response assembly in #94 but no checkpoint inference or native HTTP worker. Add the complete eager Rust/CUDA path: verify the immutable released checkpoint/config/tokenizer, run independent Qwen rows, project all 255 tied output labels on CUDA with BF16 output rounding, and serve Choice/Noul/isolated Score through the existing frontend.

One SerialScheduler admission covers the complete request and GPU head. Errors synchronize before releasing admission, retire failed execution and return unavailable health. Real warmup completes before serving. Depends on #94; the main-based diff includes that CPU contract.

Current integration and validation — 2026-10-09

Normal merge head b3871b415f0eab95b0317ffef38ef6ad276a3e86 includes main 0d2521035107e35a2670b4df716f9728a074d520 through corrected parent #94 at e0f297ee61703392d91bc26595d5da59f7385589. The workspace union retains JEV-VL and Decider; lockfile package versions remain pinned. Merge the prerequisite series in order: #94 → #110 → #111 → #112.

Fresh isolated-head checks pass: cargo fmt --all --check; cargo clippy --workspace --locked --all-targets -- -D warnings; cargo test --workspace --locked (140 passed, 0 failed, 25 ignored); cargo build --workspace --release --locked; strict MkDocs; 2 Python Decider tests and CUDA Python helper regressions. Ignored device/artifact cases are not counted as executed GPU tests.

The shared Qwen/backend now includes main's JEV-VL prefix APIs and ABI6. Rebuild worker and CUDA library together. No CUDA compilation, GPU/reference/HTTP/performance campaign ran on this macOS host; historical device results below do not establish numerical parity for the changed shared-backend integration.

Current main-based diff: 32 files, 2,481 authored source/test/build lines, 2,707 total changed lines, additions plus deletions; documentation, static fixtures, licenses and lockfiles are excluded from authored code. Core implementation/tests and required build/deployment/license inputs remain in scope; generated measurements remain external. Diagnostic helper: cargo build -p omni-decider-native --example decider-run --release --locked, executable target/release/examples/decider-run, source tests/decider/runner.rs.

Historical device evidence

2026-10-07 sequential validation archive, asset decider-next-20261007-final-evidence.tar.gz, SHA256 9f57d969fc484c518d8ebac7b2dd38174e2c6954e7af04b9352550675053ae48. This feature's executed source was 9a09e5a58d33692edd5aab5634f268055e4c8151 over main 99865743d27316fbe81362dc0f0e6de6fda86284.

2026-10-08 raw verification evidence, SHA256 5df66bf3974d534875946168fba4d7fe543d08edbd658c3756f17d6f0c8f5773. Source manifests and original protocol gates bind these results to their recorded revisions. Eager reference checks observed maximum probability/fit drift 0.0119, Score drift 0.01 and zero Choice disagreements under unchanged 0.02/0.1/0.05 gates. Packing/Graph/fault results are separately scoped in that archive; no new timing or accuracy claim is made at today's head.

Self-review

Full component and conflict-resolution review covered the cumulative CPU contract, executor/worker boundaries, tests and deployment requirements. Agent-assisted contributor review does not replace maintainer approval.

  • Reviewed the full diff and addressed findings.
  • Checked architecture ownership, dependency order and focused scope.
  • Ran applicable checks and stated ignored cases and missing GPU integration coverage.
  • Checked description, counts and numerical/performance claims against source-scoped evidence.

Exact-head GitHub checks

Head b3871b415f0eab95b0317ffef38ef6ad276a3e86: Docs passes, CI passes. Docs deploy is intentionally skipped for pull requests.

@Levius-Fubuki
Levius-Fubuki marked this pull request as ready for review October 8, 2026 01:11
Copilot AI balanced review requested due to automatic review settings October 8, 2026 01:11

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at head 5d1302aac3 (7 commits; true delta 4d4c504e..5d1302aac — the GitHub files-changed view also carries #94 because #94 sits on the older main tip). No actionable findings.

The delta adds checkpoint.rs with real identity pinning (size + SHA-256 of config/tokenizer/weights, sharded-index refusal, backbone-shape assertion, BF16 finite check, 256-row zero-padded tied-head extraction), executor.rs with a dedicated head stream and cuBLASLt GEMM handle, catch_unwind around device work, both streams synchronized before admission release, and retire-on-error setting health unavailable; serve.rs maps 415/422/500-else-503 with an 8 MiB body limit consistent with the frontend limits, and prepares requests outside admission via spawn_blocking. Error retirement and health semantics match the runtime contract pattern in docs/architecture.md, and GPU-touching unit tests are correctly #[ignore]d (fail_head.c is an LD_PRELOAD interposer for an opt-in failure test, not a CI stub).

Template, self-review checklist, and stated counts (2,481 authored / 2,707 total — both match my independent count) are in order. One P3: the body's "Full self-review completed" section still pins superseded head 31acb29b6 with 2477/3116 counts, while the Validation section correctly describes the current head — please update or drop the stale section (same staleness exists in #111–#113).

CI (rust/build/benchmarks) passes on this head (observed). The GPU campaign (RTX 4090, pinned-reference parity with max probability drift 0.0119 and zero Choice disagreements, direct+frontend HTTP parity) is author-reported at snapshot heads with SHA-pinned archives; not independently executed by this review — no GPU was used and no Rust toolchain exists on this host.

Provenance: canonical .agents/skills/system1-omni-review/SKILL.md (SHA-256 58be3bc6…dde108) and its repository map read at trusted base 4a79980d8a75, plus CONTRIBUTING.md and docs/architecture.md. Remote head rechecked before posting.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants