cua_s1: accept image embeddings and 3D positions in native language model - #56
Levius-Fubuki wants to merge 4 commits into
Conversation
|
Reviewed The HTTP text worker gives the same answers as before on this card:
Today's worker never calls
The 139-token request itself takes about 12.8 ms. On the 712-token request the extra time was smaller and noisier, and I didn't look into why. Text outputs after a multimodal call matched the text-only ones every time. Keeping a device copy of the text tables and copying it back would remove nearly all of this. #52 is on main now, so this needs a rebase, and the Minor:
|
93a6cb0 to
dea93f2
Compare
dea93f2 to
75a80cb
Compare
|
@twu3202 Rebased onto main
Added a full-model GPU regression for first-use misses, hits with changed IDs, eviction after eight cached lengths, mixed multimodal/text calls and scratch growth, comparing hidden states against eager execution and asserting capture succeeds. The existing feature-insertion regression remains. Local format, Clippy, release build, all 25 Rust CPU tests, all 7 benchmark tests, smoke-manifest validation and strict Docs build pass. Five GPU tests are ignored by default. I could not execute the new/rebased GPU checks: the server used for the original evidence is shut down and SSH refuses connections. I have therefore changed the PR to draft pending fresh ABI-3 GPU parity/replay checks. The PR description distinguishes current CPU results from the historical Update: GitHub Rust, benchmark and Docs CI all pass on |
Purpose
Add
Model::forward_multimodalfor adapted BF16 image embeddings and explicit int64 T/H/W positions. Validate unpadded inputs, insert image features into ordered placeholders, and execute interleaved MRoPE. Keep custom rotary buffers separate from immutable text positions and preserve the text graph cache.Refresh against current main's shared
omni-qwen3-5-nativemodule and 64-entry cache. Cua-S1 retains compatibility reexports; Open-Jev uses the same shared backbone. This PR changes four runtime source files. CUDA ABI remains 4, matching main.Current-main integration
Merge current main, retaining the Laya workspace and shared JSON fix for
raw_value/arbitrary_precisionfeature unification. PR64 merges the updated #56, with the sole conflict inCargo.lock. Resolve from main's lock and add only required existing image dependencies: every package/version from main and prior PR64 is retained, with no new package versions outside that union. Net runtime diff remains four files for #56 / 27 files for #64; no validation assets or unrelated documentation added.Test Plan
Baseline:
7f39ac40902c374803992407bb26eeba29c8a588.Current head:
c5eb7993d0fc313f87b8a0a91ebe9964ae77d21a.Fresh checks: formatting, strict all-target Clippy, workspace tests, locked release build and current-head CI. The frozen 11 screenshot requests / 19 questions exercise the shared JSON parser and model processor on CPU, with
arbitrary_precision,raw_value,preserve_orderandfloat_roundtripexplicitly enabled. Exact token IDs, T/H/W positions, grids, question/candidate order and usage must match. Probability/confidence drift from the previous live HTTP outputs must stay below1e-7; JSON number lexical spelling is not required to be byte-identical across Cargo feature unification.Test Result
Formatting, strict workspace Clippy, workspace tests (82 passed / 0 failed / 10 ignored) and locked release build passed locally on both current heads. Ignored tests require real checkpoints, CUDA or Laya Hopper artifacts and are not reported as CPU passes. The frozen CPU corpus passed 11 requests / 19 questions, and all three processor regressions passed invalid/missing outputs, order/usage and whole-request preparation failures.
No new standalone GPU suite was executed for #56 in this merge-only update; its numeric source is unchanged. The new-instance ABI5 vision primitive check belongs to the integrated #64 library.
The main integration changes dependency/JSON behavior; verified model/executor, CUDA bindings/kernels, image preprocessing and shared scheduler source paths are byte-identical to their previously validated revisions. No current-head full-checkpoint or live GPU HTTP run was performed in this integration update. The new instance had no Cua checkpoint; full weights were not downloaded because the changed risk is request processing and Cargo feature integration.
Earlier full-checkpoint evidence
GPU evidence belongs to PR56
ede0c05ce0f93cb74742c5feea8b7b438face1f4, basef594d7dfc4c2bef812e23f7ed73573be9625b287, not this new merge head. On RTX4090 / driver595.71.05 / CUDA13.0.88 sm89 / Rust1.99.0, six isolated-target CUDA regressions and the real-checkpoint multimodal/graph tests passed: misses/hits, 64-entry eviction, changed tokens, text→multimodal→text, and scratch growth to1,025. Matched main/#56/#64 eager/graph probes returned identical hidden values across 12 text sequences per run.The checkpoint used pinned
Qwen/Qwen3.5-4B@851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0aplus multimodalcua-ai/cua-s1-4b-0.2@16818868b0cc7813808aae4e87b417657046ab79, freshly merged in BF16 with PyTorch2.14.0+cu130 / Transformers5.17.0 / PEFT0.21.0. Download size/SHA256 checks passed. Two earlier harness/build-target failures were preserved and corrected without runtime patches or changed assertions. Full27B Open-Jev checkpoint inference and performance remain unverified; its CPU/workspace contracts passed. Integrated visual/HTTP evidence is documented separately in #64.Demo / evidence
Current integration protocol, source-equivalence proof, CPU corpus JSON, workspace/processor logs, new GPU primitive environment/binary hash and independent review are retained locally in
artifacts/cua-main-integration-20261005/; previous full-checkpoint controls and failed logs are retained inartifacts/cua-native-refresh-20261005/. They remain outside the core-code PR diff. External historical harness sources: language, vision; updated golden-corpus and graph harnesses are local, not publicly attached. No video is needed for this API/CPU parity update.Current-head GitHub Rust, benchmark and docs CI passed; deployment skipped. CI run.
Self-review
Reviewed the complete net diff, current architecture and integration delta. Independent review is recorded locally. Ready status is retained as authorized by the contributor; checklist below remains for contributor confirmation.