Repository navigation
cua_s1: add native multimodal language CUDA Graph replay - #103
Open
Levius-Fubuki wants to merge 1 commit into
Open
Levius-Fubuki wants to merge 1 commit into
Levius-Fubuki wants to merge 1 commit into
Conversation
hsliuustc0106
left a comment
Contributor
There was a problem hiding this comment.
审查了 head b418230f53a158fcbea706c1815c529a8b3832ec(base 47eff9cdeda01e4847a4fb9634a43f2cab6a233f),未发现可确认的代码缺陷。
- 输入在 replay 前重新上传;cache miss 保留 eager 结果,避免重复推进残差。文本/多模态缓存分离,扩容和捕获失败时的清理逻辑正确,视觉仍走 eager。
- 原始证据的源码 hash 与 PR 一致。重算确认 Graph ON/OFF 延迟下降 7.443%,360 组配对响应字节完全一致;该结果限于已预热、1024 tokens 的测试配置。
- 新录制日志显示文本 254 次、多模态 105 次调用均为 replay,没有新增 capture。证据包
一个维护建议:把独立 evidence 分支上的回归测试及注册另行合入主仓库,否则当前 CI 无法持续覆盖新增 graph 路径。
本地 cargo fmt --all --check、严格 workspace Clippy、workspace tests 均通过:89 passed、0 failed、12 ignored。Clippy 和 tests 使用 --locked --offline。GPU 测试核对了作者日志,未现场重跑;视频未做逐帧视觉检查。
Contributor
|
fix conflicts please |
Levius-Fubuki
force-pushed
the
codex/cua-native-mm-graph
branch
from
October 7, 2026 03:21
b418230 to
3ef72e4
Compare
Collaborator
Author
|
fixed |
4 tasks done
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Enable opt-in native multimodal language CUDA Graph replay with
CUA_S1_GRAPH=1. Token embeddings, BF16 image rows and T/H/W rotary tables are uploaded before replay. Vision remains eager. Each mode has a separate 64-entry FIFO; misses retain the eager result, and growth/drop or capture failure clears both caches. Failure diagnostics remain unconditional and capture failure disables Graph for both modes.Rebased onto main
99865743d27316fbe81362dc0f0e6de6fda86284to resolve the conflict with #102. Preserve its packed text prefill, ordered sequence-length cache keys, per-sequence math and output order. Text keys such as[4],[1,3]and[3,1]stay distinct; multimodal uses singleton[t]in its separate cache. The full upstreamrun()body remains unchanged.The core diff still contains only
src/models/qwen3_5/native/src/model.rs(63 insertions, 16 deletions). Root regressions and their wiring remain in dependent Draft #106. No kernel, ABI, dependency, weight, generated evidence or test body is added here.CUA_S1_GRAPH_TRACE=1optionally logs successful capture/replay.Test Plan
System1-Omni Version / Commit: base
99865743d27316fbe81362dc0f0e6de6fda86284; core head3ef72e4d79ce2189532a8da2373bfb2d0937273e; dependent test head2dd88db234b99a64061dbfd0094a4ab6b0b4e5cb.cargo fmt --all --check cargo clippy --workspace --locked --offline --all-targets -- -D warnings cargo test --workspace --locked --offline cargo build --workspace --release --locked --offlineUpdated root GPU tests and setup. The runner executes three ignored cases serially with tracing unset and requires both unconditional failure diagnostics. Coverage includes current token/image/position uploads, misses/replay, text/multimodal interleave, scratch growth, FIFO eviction/recapture and recording-error recovery. #106 now also checks packed text
[4]/[1,3]/[3,1], changed IDs and same-shape eager/replay outputs without aliasing equal total lengths. This does not claim bitwise batched/independent eager equivalence or cover every CUDA failure class.Test Result
Fresh local checks above all passed for the rebased core: 90 passed, 0 failed, 13 ignored. The updated dependent tests also pass all four checks: 90 passed, 0 failed, 16 ignored; the three new GPU cases compile/discover but are ignored in normal CPU execution. Complete rebase/self-review logs, source hashes, patches and preserved old-test compilation failure. The core-prefix and unchanged-upstream-run byte checks pass.
Current-head GPU execution, performance and recording are unverified. The previously supplied GPU server refused SSH connections; no paid restart or new campaign was performed. Prior head
b418230fand test head975201capassed three explicit GPU regressions and both diagnostic gates; those results are historical and do not validate the expanded cases/current rebased head. New-head GitHub Rust, benchmark and docs CI all passed (deployment skipped). Old-head CI is not carried forward as current validation.Demo / evidence
Historical recordings, complete raw results, model/adapter pins, configuration, licenses and provenance are retained at their original revisions. The latest two English 30-second videos were recorded from
b418230fafter warm capture, with genuine HTTP choices/clicks, no added waits and serial panels shown at 1x. They were not re-recorded for this conflict repair.On that historical head, the isolated same-native Graph ON/OFF/ON workload (90 frozen requests x2, 10 excluded warmups, 1024 tokens, concurrency1, RTX4090/BF16; vision eager) measured 133.1753 /144.3275 /133.9943 ms medians: 7.44% lower warmed HTTP latency, with all 360 ON/OFF response pairs byte-identical. Separate whole-framework native/reference comparisons were 2.48x text /1.51x multimodal and included earlier kernel/merged-LoRA/readout optimizations. These figures are not new-head performance claims. Capture cost, peak memory, universal speedup and agent accuracy were not established.
Self-review
Full repaired diff, #102 integration, architecture boundaries, cache lifetime/failure handling, changed test cases, artifact scope, source binding and available checks were reviewed again on 2026-10-07. No actionable source-level finding remains; unavailable GPU validation is disclosed above. The contributor explicitly confirmed these checklist items on 2026-10-06; that confirmation and Ready authorization are retained, with the new-head verification limits reported here.