Skip to content

cua_s1: add native multimodal language CUDA Graph replay - #103

Open
Levius-Fubuki wants to merge 1 commit into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-native-mm-graph
Open

Levius-Fubuki wants to merge 1 commit into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-native-mm-graph

Conversation

@Levius-Fubuki

@Levius-Fubuki Levius-Fubuki commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Enable opt-in native multimodal language CUDA Graph replay with CUA_S1_GRAPH=1. Token embeddings, BF16 image rows and T/H/W rotary tables are uploaded before replay. Vision remains eager. Each mode has a separate 64-entry FIFO; misses retain the eager result, and growth/drop or capture failure clears both caches. Failure diagnostics remain unconditional and capture failure disables Graph for both modes.

Rebased onto main 99865743d27316fbe81362dc0f0e6de6fda86284 to resolve the conflict with #102. Preserve its packed text prefill, ordered sequence-length cache keys, per-sequence math and output order. Text keys such as [4], [1,3] and [3,1] stay distinct; multimodal uses singleton [t] in its separate cache. The full upstream run() body remains unchanged.

The core diff still contains only src/models/qwen3_5/native/src/model.rs (63 insertions, 16 deletions). Root regressions and their wiring remain in dependent Draft #106. No kernel, ABI, dependency, weight, generated evidence or test body is added here. CUA_S1_GRAPH_TRACE=1 optionally logs successful capture/replay.

Test Plan

System1-Omni Version / Commit: base 99865743d27316fbe81362dc0f0e6de6fda86284; core head 3ef72e4d79ce2189532a8da2373bfb2d0937273e; dependent test head 2dd88db234b99a64061dbfd0094a4ab6b0b4e5cb.

cargo fmt --all --check
cargo clippy --workspace --locked --offline --all-targets -- -D warnings
cargo test --workspace --locked --offline
cargo build --workspace --release --locked --offline

Updated root GPU tests and setup. The runner executes three ignored cases serially with tracing unset and requires both unconditional failure diagnostics. Coverage includes current token/image/position uploads, misses/replay, text/multimodal interleave, scratch growth, FIFO eviction/recapture and recording-error recovery. #106 now also checks packed text [4] / [1,3] / [3,1], changed IDs and same-shape eager/replay outputs without aliasing equal total lengths. This does not claim bitwise batched/independent eager equivalence or cover every CUDA failure class.

Test Result

Fresh local checks above all passed for the rebased core: 90 passed, 0 failed, 13 ignored. The updated dependent tests also pass all four checks: 90 passed, 0 failed, 16 ignored; the three new GPU cases compile/discover but are ignored in normal CPU execution. Complete rebase/self-review logs, source hashes, patches and preserved old-test compilation failure. The core-prefix and unchanged-upstream-run byte checks pass.

Current-head GPU execution, performance and recording are unverified. The previously supplied GPU server refused SSH connections; no paid restart or new campaign was performed. Prior head b418230f and test head 975201ca passed three explicit GPU regressions and both diagnostic gates; those results are historical and do not validate the expanded cases/current rebased head. New-head GitHub Rust, benchmark and docs CI all passed (deployment skipped). Old-head CI is not carried forward as current validation.

Demo / evidence

Historical recordings, complete raw results, model/adapter pins, configuration, licenses and provenance are retained at their original revisions. The latest two English 30-second videos were recorded from b418230f after warm capture, with genuine HTTP choices/clicks, no added waits and serial panels shown at 1x. They were not re-recorded for this conflict repair.

On that historical head, the isolated same-native Graph ON/OFF/ON workload (90 frozen requests x2, 10 excluded warmups, 1024 tokens, concurrency1, RTX4090/BF16; vision eager) measured 133.1753 /144.3275 /133.9943 ms medians: 7.44% lower warmed HTTP latency, with all 360 ON/OFF response pairs byte-identical. Separate whole-framework native/reference comparisons were 2.48x text /1.51x multimodal and included earlier kernel/merged-LoRA/readout optimizations. These figures are not new-head performance claims. Capture cost, peak memory, universal speedup and agent accuracy were not established.

Self-review

Full repaired diff, #102 integration, architecture boundaries, cache lifetime/failure handling, changed test cases, artifact scope, source binding and available checks were reviewed again on 2026-10-07. No actionable source-level finding remains; unavailable GPU validation is disclosed above. The contributor explicitly confirmed these checklist items on 2026-10-06; that confirmation and Ready authorization are retained, with the new-head verification limits reported here.

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

审查了 head b418230f53a158fcbea706c1815c529a8b3832ec(base 47eff9cdeda01e4847a4fb9634a43f2cab6a233f),未发现可确认的代码缺陷。

  • 输入在 replay 前重新上传;cache miss 保留 eager 结果,避免重复推进残差。文本/多模态缓存分离,扩容和捕获失败时的清理逻辑正确,视觉仍走 eager。
  • 原始证据的源码 hash 与 PR 一致。重算确认 Graph ON/OFF 延迟下降 7.443%,360 组配对响应字节完全一致;该结果限于已预热、1024 tokens 的测试配置。
  • 新录制日志显示文本 254 次、多模态 105 次调用均为 replay,没有新增 capture。证据包

一个维护建议:把独立 evidence 分支上的回归测试及注册另行合入主仓库,否则当前 CI 无法持续覆盖新增 graph 路径。

本地 cargo fmt --all --check、严格 workspace Clippy、workspace tests 均通过:89 passed、0 failed、12 ignored。Clippy 和 tests 使用 --locked --offline。GPU 测试核对了作者日志,未现场重跑;视频未做逐帧视觉检查。

@Levius-Fubuki
Levius-Fubuki marked this pull request as ready for review October 6, 2026 14:52
Copilot AI balanced review requested due to automatic review settings October 6, 2026 14:52

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@hsliuustc0106

Copy link
Copy Markdown
Contributor

fix conflicts please

@Levius-Fubuki
Levius-Fubuki force-pushed the codex/cua-native-mm-graph branch from b418230 to 3ef72e4 Compare October 7, 2026 03:21
@Levius-Fubuki

Copy link
Copy Markdown
Collaborator Author

fixed

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants