Skip to content

cua_s1: complete native vision and screenshot inference with GPU parity - #64

Merged
hsliuustc0106 merged 24 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-native-vision
Oct 6, 2026
Merged

hsliuustc0106 merged 24 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-native-vision

Conversation

@Levius-Fubuki

@Levius-Fubuki Levius-Fubuki commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Complete native screenshot inference: PNG/JPEG → RGB preprocessing → 24-block CUDA vision encoder and merger → image feature insertion and T/H/W positions → language execution → candidate decisions. omni-cua-s1-vision serves the screenshot contract and reuses image features across questions in each request.

Includes #59, #63 and the refreshed #56. Vision retains BF16 base weights and all 50 FP32 LoRA pairs separately; language uses a merged BF16 export. Startup verifies pinned checkpoint/export identity and hashes, then executes real warmup before readiness.

Refresh against current main's shared Qwen module, optimized kernels and 64-entry graph cache. CPU decoding/tokenization and response reconstruction live in the processor, outside executor admission. Shared SerialScheduler admits one complete request; both CUDA streams drain before release, including failures. Multimodal execution remains eager. CUDA ABI is 5: rebuild both library and worker.

The diff retains runtime code, required dependencies, upstream licenses/notices and the one-time exporter. External tests, reports and controls are kept outside the PR diff.

Build and launch

Obtain the pinned weights/lock and Python environment using recipe/cua_s1/text.md; additionally install torchvision==0.29.0 numpy==2.5.3 Pillow==11.3.0.

PYTHONPATH=src .venv/bin/python recipe/cua_s1/export_multimodal_language.py \
  --base weights/Qwen3.5-4B --adapter weights/cua-s1-4b-0.2/multimodal \
  --out weights/cua-s1-multimodal-language
src/backends/cuda/qwen3_5/build.sh target/release 89
cargo build --release --locked -p omni-cua-s1-native --bins
CUA_S1_BASE=weights/Qwen3.5-4B \
CUA_S1_VISION_ADAPTER=weights/cua-s1-4b-0.2/multimodal \
CUA_S1_MODEL=weights/cua-s1-multimodal-language \
CUA_S1_CUDA_LIB=$PWD/target/release/libqwen3_5_cuda.so \
  target/release/omni-cua-s1-vision

The worker defaults to 127.0.0.1:8000; /health reports readiness after initialization/warmup. Export manifest hashes supply local provenance, not an external signature; keep checkpoint files immutable during execution.

Current-main integration and review fix

Current head: 066f5e38183e24a047be1fb9b4f44dddb0da54c3, based on main ae86032cba2466f45f42c2ebdfadbcfa8c30eb9f. The only merge conflict was Cargo.lock; resolution retains every package/version/checksum from main and prior PR64, with no versions outside their union. The core diff remains 27 files.

The JPEG request parser now walks bounded header segments before decoding and rejects APP2 MPF\0 containers, preventing silent first-picture selection for MPO input. Ordinary JPEG and unrelated APP2/COM metadata remain accepted. Three production-parser regressions cover an actual Pillow two-frame MPO, valid JPEG/metadata and truncation. The MPO test failed on the old parser and passes after the fix; independent review found no actionable issue.

Fresh validation on this head

  • Formatting, strict all-target workspace Clippy, locked workspace tests (89 passed, 0 failed, 12 ignored) and locked release build passed. Ignored hardware/checkpoint tests are not reported as CPU passes. Benchmark smoke validation and all eight benchmark tests passed under Python3.12. An initial run with the Mac's Python3.9 failed because the harness uses asyncio.timeout; its logs are retained.
  • Frozen CPU processing corpus: 11 requests / 19 questions, exact tokens, T/H/W positions, grids, question/candidate order and usage; reconstructed probability/confidence drift below 1e-7. Three processor-contract tests passed after setting the required tokenizer directory; the initial missing-environment failure is retained.
  • On the same RTX4090 / driver595.71.05 / CUDA13.0.88 instance, worker/frontend were rebuilt from this exact source archive. All 126 selected source hashes were verified. Numerical runtime and CUDA sources match the previous validated recording; the verified ABI5 CUDA library was reused. Pinned base, multimodal adapter and BF16 language export are unchanged.
  • Fresh live HTTP through both the native worker and Rust frontend: 22 successful screenshot requests / 38 questions matched frozen historical native replay outputs within 1e-7; 36 invalid-input checks passed. The actual two-frame MPO returned 422 at both endpoints, and an ordinary JPEG returned 200 at both endpoints. This is current-head HTTP regression evidence against frozen native outputs; it is not a new independent FP32/BF16 oracle evaluation.

Earlier full-checkpoint evidence

Full-checkpoint evidence is tied to PR56 ede0c05ce0f93cb74742c5feea8b7b438face1f4 / PR64 1152531a704044c879adabf377986d14010edbdb, base f594d7dfc4c2bef812e23f7ed73573be9625b287, not these new merge heads. The PR64 build source 9538f00 has the identical tree to 1152531.

One RTX4090, driver595.71.05, CUDA13.0.88/sm89, Rust1.99.0, Python3.12.3, PyTorch2.14.0+cu130, Transformers5.17.0, PEFT0.21.0. Pinned Qwen/Qwen3.5-4B@851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a and cua-ai/cua-s1-4b-0.2@16818868b0cc7813808aae4e87b417657046ab79; downloaded size/SHA256 checks passed. Vision used BF16 base / separate FP32 LoRA; merged language BF16; independent FP32/BF16 controls, TF32 off.

Six isolated-target CUDA regressions per revision, full-model graph misses/hits/64-entry eviction/changed inputs/scratch growth to1,025 and multimodal boundary tests passed. Main/#56/#64 eager/graph matched 12 text sequences per run. Integrated screenshot standard cases matched 8/8 choices (max probability error 0.00331324, allowance 0.02052773); boundary cases 11/11 (max 0.07950398, allowance 0.20292068). Worker and frontend each matched direct results for 11 requests/19 questions within 1e-7, with nine invalid inputs per corpus.22 requests at concurrency8 returned bodies identical to serial on that runtime. Tokens/grids/positions matched and native replay was identical; PNG decoding matched Pillow, tested JPEG pixel differences ≤3.

The rule was frozen per corpus: max native probability error ≤2×max BF16-reference error+.01, matching choices for FP32 margin≥.05. Corpus/tolerances were unchanged. Two prior PR56 harness/build-target failures are retained; they were corrected using isolated targets and explicit PR56 binaries without runtime patches or changed assertions. The full27B Open-Jev checkpoint was not run; its CPU/workspace contracts passed. These finite corpora do not prove bitwise visual equivalence or general accuracy. Latency/throughput/cold-vs-warm performance were not evaluated; no speed or memory improvement claim.

Integration demo / evidence

A fresh 62-second English video on this exact repaired head records service launch, GPU warmup, screenshot/page state, the real system1-agents DecisionModel request through Rust to the native worker, returned probabilities and actual browser clicks. The model selected Enable then Save; the final DOM state is enabled=true/saved=true. Request and response hashes match across both HTTP boundaries. Visible cursor/click highlights are cleared before model screenshots.

This uses a small external DecisionModel adapter and a Playwright loop over a controlled page. It demonstrates that integration path; it is not a built-in stock s1a CLI feature or a general agent/performance benchmark. The video and full records remain local pending attachment; only the newest recording is retained.

Fresh repair/CPU/HTTP/review logs are retained in artifacts/cua-pr64-repair-20261006/; the latest video, requests/replies, binary/source hashes and scripts are in artifacts/cua-full-recording-latest/. Historical non-video validation and failed-run logs remain outside the core PR diff.

Current-head CI status is reported by GitHub checks. No speed, memory improvement or full Open-Jev checkpoint claim is made.

Self-review

Reviewed the net diff, bounded parser, lock resolution and evidence scopes. Independent review of the repair found no actionable issue.

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description and validation claims match the implementation and available evidence.

@Levius-Fubuki
Levius-Fubuki marked this pull request as ready for review October 5, 2026 03:16
Copilot AI balanced review requested due to automatic review settings October 5, 2026 03:16

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed 55dcf0ab3da133add4efecb856d4cbcf8e343c67 against merge base 7f39ac40902c374803992407bb26eeba29c8a588. One reproduced P2 request-contract defect is attached inline.

Validation: workspace tests passed (82 passed, 0 failed, 10 ignored), formatting and strict all-target workspace Clippy passed. CPU decoder probes matched Pillow RGB bytes for 13 PNG variants; RGB/L/CMYK JPEG probes differed by at most 2/1/1 pixel levels. The two-frame MPO rejection mismatch was reproduced through the actual native parser and existing Python parser. No GPU or live worker inference was run; this does not establish current-head model parity.

GitHub also currently reports a merge conflict for this PR.

Comment on lines +115 to +118
} else {
let mut reader = image::ImageReader::with_format(Cursor::new(&raw), format);
reader.limits(limits);
reader.decode()?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Reject multi-picture JPEG inputs before decoding

This branch has no equivalent of the PNG single-frame check. A 32×32 two-frame MPO generated by Pillow and supplied as data:image/jpeg;base64,... is accepted by the production native parser, which returns only its first frame. The existing Python screenshot parser rejects the same request. Because the HTTP preparation path calls this parser, unsupported multi-picture input can proceed to inference with a silently selected screenshot instead of being rejected.

Reproduce by saving one Pillow RGB image with format="MPO", save_all=True, append_images=[second_image], base64-encoding it under the JPEG data-URL prefix, and passing the same valid choice request to both parsers. This mismatch was reproduced on CPU; no live GPU HTTP result is claimed.

Detect and reject MPF/MPO multi-picture JPEG containers before decoding, and add a regression through the production request parser.

@hsliuustc0106
hsliuustc0106 merged commit 47eff9c into ThinkFlowLab:main Oct 6, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants