cua_s1: complete native vision and screenshot inference with GPU parity - #64
Conversation
# Conflicts: # recipe/cua_s1/native.md # src/models/cua_s1/native/Cargo.toml
…e-vision # Conflicts: # recipe/cua_s1/native.md # src/models/cua_s1/native/src/lib.rs
…a-vision-loader # Conflicts: # src/models/cua_s1/native/Cargo.toml # src/models/cua_s1/native/src/lib.rs
…a-image-preprocess # Conflicts: # src/models/cua_s1/native/src/lib.rs
hsliuustc0106
left a comment
There was a problem hiding this comment.
Reviewed 55dcf0ab3da133add4efecb856d4cbcf8e343c67 against merge base 7f39ac40902c374803992407bb26eeba29c8a588. One reproduced P2 request-contract defect is attached inline.
Validation: workspace tests passed (82 passed, 0 failed, 10 ignored), formatting and strict all-target workspace Clippy passed. CPU decoder probes matched Pillow RGB bytes for 13 PNG variants; RGB/L/CMYK JPEG probes differed by at most 2/1/1 pixel levels. The two-frame MPO rejection mismatch was reproduced through the actual native parser and existing Python parser. No GPU or live worker inference was run; this does not establish current-head model parity.
GitHub also currently reports a merge conflict for this PR.
| } else { | ||
| let mut reader = image::ImageReader::with_format(Cursor::new(&raw), format); | ||
| reader.limits(limits); | ||
| reader.decode()? |
There was a problem hiding this comment.
[P2] Reject multi-picture JPEG inputs before decoding
This branch has no equivalent of the PNG single-frame check. A 32×32 two-frame MPO generated by Pillow and supplied as data:image/jpeg;base64,... is accepted by the production native parser, which returns only its first frame. The existing Python screenshot parser rejects the same request. Because the HTTP preparation path calls this parser, unsupported multi-picture input can proceed to inference with a silently selected screenshot instead of being rejected.
Reproduce by saving one Pillow RGB image with format="MPO", save_all=True, append_images=[second_image], base64-encoding it under the JPEG data-URL prefix, and passing the same valid choice request to both parsers. This mismatch was reproduced on CPU; no live GPU HTTP result is claimed.
Detect and reject MPF/MPO multi-picture JPEG containers before decoding, and add a regression through the production request parser.
Purpose
Complete native screenshot inference: PNG/JPEG → RGB preprocessing → 24-block CUDA vision encoder and merger → image feature insertion and T/H/W positions → language execution → candidate decisions.
omni-cua-s1-visionserves the screenshot contract and reuses image features across questions in each request.Includes #59, #63 and the refreshed #56. Vision retains BF16 base weights and all 50 FP32 LoRA pairs separately; language uses a merged BF16 export. Startup verifies pinned checkpoint/export identity and hashes, then executes real warmup before readiness.
Refresh against current main's shared Qwen module, optimized kernels and 64-entry graph cache. CPU decoding/tokenization and response reconstruction live in the processor, outside executor admission. Shared
SerialScheduleradmits one complete request; both CUDA streams drain before release, including failures. Multimodal execution remains eager. CUDA ABI is 5: rebuild both library and worker.The diff retains runtime code, required dependencies, upstream licenses/notices and the one-time exporter. External tests, reports and controls are kept outside the PR diff.
Build and launch
Obtain the pinned weights/lock and Python environment using
recipe/cua_s1/text.md; additionally installtorchvision==0.29.0 numpy==2.5.3 Pillow==11.3.0.PYTHONPATH=src .venv/bin/python recipe/cua_s1/export_multimodal_language.py \ --base weights/Qwen3.5-4B --adapter weights/cua-s1-4b-0.2/multimodal \ --out weights/cua-s1-multimodal-language src/backends/cuda/qwen3_5/build.sh target/release 89 cargo build --release --locked -p omni-cua-s1-native --bins CUA_S1_BASE=weights/Qwen3.5-4B \ CUA_S1_VISION_ADAPTER=weights/cua-s1-4b-0.2/multimodal \ CUA_S1_MODEL=weights/cua-s1-multimodal-language \ CUA_S1_CUDA_LIB=$PWD/target/release/libqwen3_5_cuda.so \ target/release/omni-cua-s1-visionThe worker defaults to
127.0.0.1:8000;/healthreports readiness after initialization/warmup. Export manifest hashes supply local provenance, not an external signature; keep checkpoint files immutable during execution.Current-main integration and review fix
Current head:
066f5e38183e24a047be1fb9b4f44dddb0da54c3, based on mainae86032cba2466f45f42c2ebdfadbcfa8c30eb9f. The only merge conflict wasCargo.lock; resolution retains every package/version/checksum from main and prior PR64, with no versions outside their union. The core diff remains 27 files.The JPEG request parser now walks bounded header segments before decoding and rejects APP2
MPF\0containers, preventing silent first-picture selection for MPO input. Ordinary JPEG and unrelated APP2/COM metadata remain accepted. Three production-parser regressions cover an actual Pillow two-frame MPO, valid JPEG/metadata and truncation. The MPO test failed on the old parser and passes after the fix; independent review found no actionable issue.Fresh validation on this head
asyncio.timeout; its logs are retained.1e-7. Three processor-contract tests passed after setting the required tokenizer directory; the initial missing-environment failure is retained.1e-7; 36 invalid-input checks passed. The actual two-frame MPO returned 422 at both endpoints, and an ordinary JPEG returned 200 at both endpoints. This is current-head HTTP regression evidence against frozen native outputs; it is not a new independent FP32/BF16 oracle evaluation.Earlier full-checkpoint evidence
Full-checkpoint evidence is tied to PR56
ede0c05ce0f93cb74742c5feea8b7b438face1f4/ PR641152531a704044c879adabf377986d14010edbdb, basef594d7dfc4c2bef812e23f7ed73573be9625b287, not these new merge heads. The PR64 build source9538f00has the identical tree to1152531.One RTX4090, driver595.71.05, CUDA13.0.88/sm89, Rust1.99.0, Python3.12.3, PyTorch2.14.0+cu130, Transformers5.17.0, PEFT0.21.0. Pinned
Qwen/Qwen3.5-4B@851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0aandcua-ai/cua-s1-4b-0.2@16818868b0cc7813808aae4e87b417657046ab79; downloaded size/SHA256 checks passed. Vision used BF16 base / separate FP32 LoRA; merged language BF16; independent FP32/BF16 controls, TF32 off.Six isolated-target CUDA regressions per revision, full-model graph misses/hits/64-entry eviction/changed inputs/scratch growth to1,025 and multimodal boundary tests passed. Main/#56/#64 eager/graph matched 12 text sequences per run. Integrated screenshot standard cases matched 8/8 choices (max probability error 0.00331324, allowance 0.02052773); boundary cases 11/11 (max 0.07950398, allowance 0.20292068). Worker and frontend each matched direct results for 11 requests/19 questions within
1e-7, with nine invalid inputs per corpus.22 requests at concurrency8 returned bodies identical to serial on that runtime. Tokens/grids/positions matched and native replay was identical; PNG decoding matched Pillow, tested JPEG pixel differences ≤3.The rule was frozen per corpus: max native probability error ≤2×max BF16-reference error+.01, matching choices for FP32 margin≥.05. Corpus/tolerances were unchanged. Two prior PR56 harness/build-target failures are retained; they were corrected using isolated targets and explicit PR56 binaries without runtime patches or changed assertions. The full27B Open-Jev checkpoint was not run; its CPU/workspace contracts passed. These finite corpora do not prove bitwise visual equivalence or general accuracy. Latency/throughput/cold-vs-warm performance were not evaluated; no speed or memory improvement claim.
Integration demo / evidence
A fresh 62-second English video on this exact repaired head records service launch, GPU warmup, screenshot/page state, the real system1-agents DecisionModel request through Rust to the native worker, returned probabilities and actual browser clicks. The model selected Enable then Save; the final DOM state is enabled=true/saved=true. Request and response hashes match across both HTTP boundaries. Visible cursor/click highlights are cleared before model screenshots.
This uses a small external DecisionModel adapter and a Playwright loop over a controlled page. It demonstrates that integration path; it is not a built-in stock s1a CLI feature or a general agent/performance benchmark. The video and full records remain local pending attachment; only the newest recording is retained.
Fresh repair/CPU/HTTP/review logs are retained in
artifacts/cua-pr64-repair-20261006/; the latest video, requests/replies, binary/source hashes and scripts are inartifacts/cua-full-recording-latest/. Historical non-video validation and failed-run logs remain outside the core PR diff.Current-head CI status is reported by GitHub checks. No speed, memory improvement or full Open-Jev checkpoint claim is made.
Self-review
Reviewed the net diff, bounded parser, lock resolution and evidence scopes. Independent review of the repair found no actionable issue.