Skip to content

cua_s1: add native vision CUDA Graph replay - #108

Open
Levius-Fubuki wants to merge 1 commit into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-native-vision-graph
Open

Levius-Fubuki wants to merge 1 commit into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-native-vision-graph

Conversation

@Levius-Fubuki

@Levius-Fubuki Levius-Fubuki commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Native screenshot requests currently reallocate vision activations and launch every vision operator eagerly. Retain one exact [T,H,W] scratch set and add opt-in full vision encoder CUDA Graph replay with CUA_S1_VISION_GRAPH=1. The capture includes the 24 vision blocks, unmerged FP32 LoRA branches and merger; pixel upload, CPU processing and synchronized feature download remain outside it.

Eager and Graph share the retained scratch control. A grid change synchronizes and retires the old executable before releasing its buffers; equal patch counts with different geometry cannot reuse a capture. Every replay uploads the new pixels. A miss returns the completed eager result after capture; capture failure logs unconditionally, preserves that result and disables vision Graph for the loaded model. Trace runs eagerly; launch errors propagate. Graph and scratch retire before weights, GEMM workspace and stream.

Independent of the language Graph switch and #103/#106. No CUDA kernel, ABI, precision, checkpoint, prompt, batching or transport change. Tests/helpers stay under root tests/; source contains only private-test module wiring. 503 changed authored-code lines; 606 total diff lines, across nine files. No generated measurements or media are committed.

Test Plan

System1-Omni Version / Commit: baseline 99865743d27316fbe81362dc0f0e6de6fda86284; head 43214183d41c44cc1cee50576d2458bd9661e68e.

Build, launch, GPU regression and native reproduction instructions. Use pinned Qwen3.5-4B 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, multimodal adapter 16818868b0cc7813808aae4e87b417657046ab79 and a verified BF16 language export. One RTX 4090; CUDA library ABI5, existing kernel sources matching current backend code. Exact driver/toolchain/library/binary/source hashes are in the evidence.

  • cargo fmt --all --check
  • cargo clippy --workspace --locked --all-targets -- -D warnings
  • cargo test --workspace --locked
  • cargo build --workspace --release --locked
  • Explicit release GPU tests: normal vision_graph_ regressions and isolated vision_capture_failure_ using the documented test-only wrapper. The wrapper delegates real math, ends actual capture, destroys its executable, then injects an instantiation error. It tests the actual forward's completed eager result, permanent disablement and changed-grid fallback.
  • Two measured repetitions with reversed mode order: original main, retained-scratch eager, vision Graph. Native: 10 changed screenshots at each of 256x256, 512x256 and 512x512, two questions each. HTTP: 10 synthetic 256x256 screenshots, two questions each. Concurrency 1. One feasibility configuration and warmups excluded; loading/cold times preserved separately. Exact BF16 feature digests and complete response equality required.

Test Result

Observed GitHub CI on this exact head: Rust and benchmark jobs passed; strict Docs build passed. Docs deployment is intentionally skipped for pull requests. GitHub reports no merge conflicts.

Four Rust checks pass: 90 passed / 0 failed / 16 ignored in the ordinary workspace run. Final explicit GPU checks 3/3 pass. A disabled-capture mutant correctly fails at missing vision capture; trace-disabled failure diagnostics are verified. Strict MkDocs, benchmark smoke validation and Python 3.12 benchmark tests (8/8) pass. An initial supplemental Python 3.9 run fails on existing asyncio.timeout; its log and the supported-version rerun are retained, without changing source or acceptance tolerance.

180 native timed samples and 60 direct-worker HTTP requests, with identical BF16 feature digests and complete native responses across six configurations; all HTTP bodies match byte-for-byte. No measurement failures. Production VisionModel and benchmark source are byte-identical to the measured versions; review followups only strengthen root tests/fixture and documentation.

Mean milliseconds, run 1 / run 2:

Timing boundary Original main Vision Graph Reduction vs main
Vision validation/conversion/upload/forward/download 10.741 / 10.754 7.991 / 8.021 25.60% / 25.42%
Complete prepared native decision, two questions 57.701 / 57.418 54.748 / 54.813 5.12% / 4.54%
Direct-worker HTTP, including decode/preparation/JSON 44.723 / 44.825 43.308 / 43.245 3.16% / 3.52%

Most benefit comes from scratch/geometry reuse. Graph's additional mean reduction versus retained-scratch eager is 3.52–4.15% vision, 0.46–0.51% complete native decision and 0.76–1.51% HTTP. First-use capture adds cost: 512x256 first vision is 11.71–11.75 ms versus main's 9.44–9.57 ms. Not every p95 improves versus retained-scratch eager. At the fixed 256x256 HTTP workload, sampled device resident memory is main 9386 MiB, eager 9402 MiB, Graph 9406 MiB. Native memory polls every 200 ms can miss transient peaks.

Serial synthetic measurements on one GPU, two repetitions; no confidence-interval, labelled task-quality, production/concurrency throughput, frontend-proxy speed, long-term cache-churn, other-GPU, Python or Metal claim. Native decision timing excludes CPU preparation and HTTP; model loading and warmup are excluded from warm comparisons. Language remains eager in all measured multimodal modes.

Demo / evidence

Raw results, complete synthetic inputs/responses, hashes, failures, protocol and self-review archive · summary JSON · self-review report.

Archive SHA256: 46ecdb36ebe5f5dc77bbe0cbebab5aa6949f8a7bfd6b89b3cd825a84e9b987e8. Screenshots are deterministic task-created RGB patterns; no private screens, credentials or model weights are published. Video N/A: the changed behavior is capture/replay and feature parity, shown by raw measurements and trace logs.

Self-review

Full agent precheck and independent code review completed; two test/documentation findings were fixed and rechecked, with no remaining blocker. Self-review was refreshed against main 99865743d27316fbe81362dc0f0e6de6fda86284 on 2026-10-07, with no remaining blocker. The author explicitly authorized Ready conversion after self-review. All four Rust checks were rerun successfully on the unchanged head; GPU/parity evidence and current CI cover that same head. Agent-assisted self-review does not replace maintainer review.

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

@Levius-Fubuki
Levius-Fubuki marked this pull request as ready for review October 7, 2026 05:44
Copilot AI balanced review requested due to automatic review settings October 7, 2026 05:44

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants