Repository navigation
cua_s1: add native vision CUDA Graph replay - #108
Open
Levius-Fubuki wants to merge 1 commit into
Open
Levius-Fubuki wants to merge 1 commit into
Levius-Fubuki wants to merge 1 commit into
Conversation
Levius-Fubuki
marked this pull request as ready for review
October 7, 2026 05:44
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Native screenshot requests currently reallocate vision activations and launch every vision operator eagerly. Retain one exact
[T,H,W]scratch set and add opt-in full vision encoder CUDA Graph replay withCUA_S1_VISION_GRAPH=1. The capture includes the 24 vision blocks, unmerged FP32 LoRA branches and merger; pixel upload, CPU processing and synchronized feature download remain outside it.Eager and Graph share the retained scratch control. A grid change synchronizes and retires the old executable before releasing its buffers; equal patch counts with different geometry cannot reuse a capture. Every replay uploads the new pixels. A miss returns the completed eager result after capture; capture failure logs unconditionally, preserves that result and disables vision Graph for the loaded model. Trace runs eagerly; launch errors propagate. Graph and scratch retire before weights, GEMM workspace and stream.
Independent of the language Graph switch and #103/#106. No CUDA kernel, ABI, precision, checkpoint, prompt, batching or transport change. Tests/helpers stay under root
tests/; source contains only private-test module wiring. 503 changed authored-code lines; 606 total diff lines, across nine files. No generated measurements or media are committed.Test Plan
System1-Omni Version / Commit: baseline
99865743d27316fbe81362dc0f0e6de6fda86284; head43214183d41c44cc1cee50576d2458bd9661e68e.Build, launch, GPU regression and native reproduction instructions. Use pinned Qwen3.5-4B
851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, multimodal adapter16818868b0cc7813808aae4e87b417657046ab79and a verified BF16 language export. One RTX 4090; CUDA library ABI5, existing kernel sources matching current backend code. Exact driver/toolchain/library/binary/source hashes are in the evidence.cargo fmt --all --checkcargo clippy --workspace --locked --all-targets -- -D warningscargo test --workspace --lockedcargo build --workspace --release --lockedvision_graph_regressions and isolatedvision_capture_failure_using the documented test-only wrapper. The wrapper delegates real math, ends actual capture, destroys its executable, then injects an instantiation error. It tests the actual forward's completed eager result, permanent disablement and changed-grid fallback.Test Result
Observed GitHub CI on this exact head: Rust and benchmark jobs passed; strict Docs build passed. Docs deployment is intentionally skipped for pull requests. GitHub reports no merge conflicts.
Four Rust checks pass: 90 passed / 0 failed / 16 ignored in the ordinary workspace run. Final explicit GPU checks 3/3 pass. A disabled-capture mutant correctly fails at
missing vision capture; trace-disabled failure diagnostics are verified. Strict MkDocs, benchmark smoke validation and Python 3.12 benchmark tests (8/8) pass. An initial supplemental Python 3.9 run fails on existingasyncio.timeout; its log and the supported-version rerun are retained, without changing source or acceptance tolerance.180 native timed samples and 60 direct-worker HTTP requests, with identical BF16 feature digests and complete native responses across six configurations; all HTTP bodies match byte-for-byte. No measurement failures. Production
VisionModeland benchmark source are byte-identical to the measured versions; review followups only strengthen root tests/fixture and documentation.Mean milliseconds, run 1 / run 2:
Most benefit comes from scratch/geometry reuse. Graph's additional mean reduction versus retained-scratch eager is 3.52–4.15% vision, 0.46–0.51% complete native decision and 0.76–1.51% HTTP. First-use capture adds cost: 512x256 first vision is 11.71–11.75 ms versus main's 9.44–9.57 ms. Not every p95 improves versus retained-scratch eager. At the fixed 256x256 HTTP workload, sampled device resident memory is main 9386 MiB, eager 9402 MiB, Graph 9406 MiB. Native memory polls every 200 ms can miss transient peaks.
Serial synthetic measurements on one GPU, two repetitions; no confidence-interval, labelled task-quality, production/concurrency throughput, frontend-proxy speed, long-term cache-churn, other-GPU, Python or Metal claim. Native decision timing excludes CPU preparation and HTTP; model loading and warmup are excluded from warm comparisons. Language remains eager in all measured multimodal modes.
Demo / evidence
Raw results, complete synthetic inputs/responses, hashes, failures, protocol and self-review archive · summary JSON · self-review report.
Archive SHA256:
46ecdb36ebe5f5dc77bbe0cbebab5aa6949f8a7bfd6b89b3cd825a84e9b987e8. Screenshots are deterministic task-created RGB patterns; no private screens, credentials or model weights are published. Video N/A: the changed behavior is capture/replay and feature parity, shown by raw measurements and trace logs.Self-review
Full agent precheck and independent code review completed; two test/documentation findings were fixed and rechecked, with no remaining blocker. Self-review was refreshed against main
99865743d27316fbe81362dc0f0e6de6fda86284on 2026-10-07, with no remaining blocker. The author explicitly authorized Ready conversion after self-review. All four Rust checks were rerun successfully on the unchanged head; GPU/parity evidence and current CI cover that same head. Agent-assisted self-review does not replace maintainer review.