Skip to content

cua_s1: accept image embeddings and 3D positions in native language model - #56

Open
Levius-Fubuki wants to merge 4 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-native-multimodal-input
Open

Levius-Fubuki wants to merge 4 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-native-multimodal-input

Conversation

@Levius-Fubuki

@Levius-Fubuki Levius-Fubuki commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Add Model::forward_multimodal for adapted BF16 image embeddings and explicit int64 T/H/W positions. Validate unpadded inputs, insert image features into ordered placeholders, and execute interleaved MRoPE. Keep custom rotary buffers separate from immutable text positions and preserve the text graph cache.

Refresh against current main's shared omni-qwen3-5-native module and 64-entry cache. Cua-S1 retains compatibility reexports; Open-Jev uses the same shared backbone. This PR changes four runtime source files. CUDA ABI remains 4, matching main.

Current-main integration

Merge current main, retaining the Laya workspace and shared JSON fix for raw_value / arbitrary_precision feature unification. PR64 merges the updated #56, with the sole conflict in Cargo.lock. Resolve from main's lock and add only required existing image dependencies: every package/version from main and prior PR64 is retained, with no new package versions outside that union. Net runtime diff remains four files for #56 / 27 files for #64; no validation assets or unrelated documentation added.

Test Plan

Baseline: 7f39ac40902c374803992407bb26eeba29c8a588.
Current head: c5eb7993d0fc313f87b8a0a91ebe9964ae77d21a.

Fresh checks: formatting, strict all-target Clippy, workspace tests, locked release build and current-head CI. The frozen 11 screenshot requests / 19 questions exercise the shared JSON parser and model processor on CPU, with arbitrary_precision, raw_value, preserve_order and float_roundtrip explicitly enabled. Exact token IDs, T/H/W positions, grids, question/candidate order and usage must match. Probability/confidence drift from the previous live HTTP outputs must stay below 1e-7; JSON number lexical spelling is not required to be byte-identical across Cargo feature unification.

Test Result

Formatting, strict workspace Clippy, workspace tests (82 passed / 0 failed / 10 ignored) and locked release build passed locally on both current heads. Ignored tests require real checkpoints, CUDA or Laya Hopper artifacts and are not reported as CPU passes. The frozen CPU corpus passed 11 requests / 19 questions, and all three processor regressions passed invalid/missing outputs, order/usage and whole-request preparation failures.

No new standalone GPU suite was executed for #56 in this merge-only update; its numeric source is unchanged. The new-instance ABI5 vision primitive check belongs to the integrated #64 library.

The main integration changes dependency/JSON behavior; verified model/executor, CUDA bindings/kernels, image preprocessing and shared scheduler source paths are byte-identical to their previously validated revisions. No current-head full-checkpoint or live GPU HTTP run was performed in this integration update. The new instance had no Cua checkpoint; full weights were not downloaded because the changed risk is request processing and Cargo feature integration.

Earlier full-checkpoint evidence

GPU evidence belongs to PR56 ede0c05ce0f93cb74742c5feea8b7b438face1f4, base f594d7dfc4c2bef812e23f7ed73573be9625b287, not this new merge head. On RTX4090 / driver595.71.05 / CUDA13.0.88 sm89 / Rust1.99.0, six isolated-target CUDA regressions and the real-checkpoint multimodal/graph tests passed: misses/hits, 64-entry eviction, changed tokens, text→multimodal→text, and scratch growth to1,025. Matched main/#56/#64 eager/graph probes returned identical hidden values across 12 text sequences per run.

The checkpoint used pinned Qwen/Qwen3.5-4B@851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a plus multimodal cua-ai/cua-s1-4b-0.2@16818868b0cc7813808aae4e87b417657046ab79, freshly merged in BF16 with PyTorch2.14.0+cu130 / Transformers5.17.0 / PEFT0.21.0. Download size/SHA256 checks passed. Two earlier harness/build-target failures were preserved and corrected without runtime patches or changed assertions. Full27B Open-Jev checkpoint inference and performance remain unverified; its CPU/workspace contracts passed. Integrated visual/HTTP evidence is documented separately in #64.

Demo / evidence

Current integration protocol, source-equivalence proof, CPU corpus JSON, workspace/processor logs, new GPU primitive environment/binary hash and independent review are retained locally in artifacts/cua-main-integration-20261005/; previous full-checkpoint controls and failed logs are retained in artifacts/cua-native-refresh-20261005/. They remain outside the core-code PR diff. External historical harness sources: language, vision; updated golden-corpus and graph harnesses are local, not publicly attached. No video is needed for this API/CPU parity update.

Current-head GitHub Rust, benchmark and docs CI passed; deployment skipped. CI run.

Self-review

Reviewed the complete net diff, current architecture and integration delta. Independent review is recorded locally. Ready status is retained as authorized by the contributor; checklist below remains for contributor confirmation.

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

Copilot AI balanced review requested due to automatic review settings October 1, 2026 02:04

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@twu3202

twu3202 commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Reviewed 93a6cb0 on an RTX 6000 Ada (sm_89) with CUDA 13.2 and driver 595.91.07, against main 566dec1. Both builds used the same libqwen3_5_cuda.so, since the kernels are unchanged.

The HTTP text worker gives the same answers as before on this card:

  • I sent 567 requests to a main worker and a cua_s1: accept image embeddings and 3D positions in native language model #56 worker side by side: 560 text requests covering answers and validation errors, plus 7 HTTP edge cases. Status, content type and body are byte-identical on all 567, including the 134 answered ones.
  • 276 prompts (272 lengths from 173 to 2,024 tokens, plus 2,047, 2,048, 2,049 and 16,384), cold and warm, then 8 concurrent clients, four runs per build: every answer is byte-identical across builds and runs, and so is the 413 for 16,385 tokens.
  • Latency is the same to within about 1%. The spread between runs followed the card's temperature when a run started, not the build: runs starting at 55 to 58 C averaged 93.0 to 94.6 ms per warm request for both builds, and runs starting at 72 to 75 C averaged 100.4 to 102.5 ms.
  • The CPU tests and the three GPU tests pass here. For the multimodal GPU test I used our merged text weights with the base model's config.json, which adds image_token_id, because the native loader takes BF16 only and the base checkpoint stores A_log and the linear-attention norms in F32.

Today's worker never calls forward_multimodal, so the following only matters for a process that serves both. The first text call after a multimodal call rebuilds the rotary tables on the host for every scratch row, and the scratch keeps the size of the longest prompt so far. Medians of 30, with each timed text call preceded by a 2-token call (an image-free multimodal one, or a text one for the baseline):

Scratch rows Host rebuild alone Extra on a 139-token request Extra on a 712-token request
1,024 0.45 ms 0.51 ms 0.05 ms
2,048 0.94 ms 1.08 ms 0.11 ms
4,096 1.95 ms 2.19 ms 0.63 ms
8,192 3.96 ms 4.45 ms 1.73 ms
16,384 8.00 ms 8.72 ms 2.26 ms

The 139-token request itself takes about 12.8 ms. On the 712-token request the extra time was smaller and noisier, and I didn't look into why. Text outputs after a multimodal call matched the text-only ones every time. Keeping a device copy of the text tables and copying it back would remove nearly all of this.

#52 is on main now, so this needs a rebase, and the forward conflict has a trap. After this PR, run() no longer embeds, and the layers update s.res in place (add_rms_norm_kernel writes residual[i]). On a cache miss, main now runs run() eagerly and then launches the captured graph, which works only because main's run() embeds from s.ids first. If the conflict is resolved by calling embed_tokens before that block, the replay starts from the residual the eager pass already advanced, so with CUA_S1_GRAPH=1 every miss (a new length, or one evicted from the 8 entries) returns a wrong hidden state without an error. Using the eager result on a miss, or embedding again before the launch, avoids that. The numbers above are from 93a6cb0 against 566dec1, before #52 landed.

Minor:

  • The interleave unit test uses sections [12, 10, 10], so the real [11, 11, 10] layout, where index 31 belongs to H, is not covered.
  • The "Replay a reference boundary" subsection and the example depend on the bundle format from Add Cua-S1 multimodal reference tensor exports #53, which is still open.

@Levius-Fubuki
Levius-Fubuki marked this pull request as draft October 1, 2026 09:09
@Levius-Fubuki
Levius-Fubuki force-pushed the codex/cua-native-multimodal-input branch from 93a6cb0 to dea93f2 Compare October 1, 2026 09:09
@Levius-Fubuki
Levius-Fubuki force-pushed the codex/cua-native-multimodal-input branch from dea93f2 to 75a80cb Compare October 1, 2026 09:13
@Levius-Fubuki

Levius-Fubuki commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

@twu3202 Rebased onto main fd0d87e in 75a80cb and addressed the actionable items:

  • Graph misses return the warm-up eager result, then retain the captured layers for future hits. Hits first re-embed the uploaded token IDs. The captured graph excludes embedding, so the eager residual is never replayed on a miss. Existing failure fallback and cache invalidation before scratch growth/drop are retained.
  • Text rotary tables stay immutable on the device. Explicit-position calls use a separate pair of tables, so switching back to text does not rebuild host tables or copy tables back. This adds 2 MiB at 16,384 scratch rows; no CUDA backend/ABI change beyond main's ABI 3 is needed. I have not measured new latency.
  • The distinct-axis interleave oracle now covers actual [11,11,10], including H at frequency 31, in addition to the synthetic layout.
  • The recipe/example explicitly identify Add Cua-S1 multimodal reference tensor exports #53 as still open and pin its exporter revision. Only the optional replay fixture producer/verifier depend on that PR; the Rust model API can accept features/positions directly.

Added a full-model GPU regression for first-use misses, hits with changed IDs, eviction after eight cached lengths, mixed multimodal/text calls and scratch growth, comparing hidden states against eager execution and asserting capture succeeds. The existing feature-insertion regression remains.

Local format, Clippy, release build, all 25 Rust CPU tests, all 7 benchmark tests, smoke-manifest validation and strict Docs build pass. Five GPU tests are ignored by default. I could not execute the new/rebased GPU checks: the server used for the original evidence is shut down and SSH refuses connections. I have therefore changed the PR to draft pending fresh ABI-3 GPU parity/replay checks. The PR description distinguishes current CPU results from the historical 93a6cb0 GPU evidence; none of the original accuracy or latency numbers are relabelled as results for 75a80cb.

Update: GitHub Rust, benchmark and Docs CI all pass on 75a80cb. The five hardware tests remain unexecuted and the draft/GPU-validation status above is unchanged.

@Levius-Fubuki
Levius-Fubuki marked this pull request as ready for review October 5, 2026 03:15

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants