cua_s1: preprocess RGB images for native vision on CPU - #63
Merged
hsliuustc0106 merged 4 commits intoOct 6, 2026
Merged
Conversation
4 tasks done
This was referenced Oct 5, 2026
…a-image-preprocess # Conflicts: # src/models/cua_s1/native/src/lib.rs
Levius-Fubuki
marked this pull request as ready for review
October 5, 2026 15:51
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Add
image_preprocess::preprocess_rgb8(width, height, rgb)for decoded interleaved RGB8 input. Return float32 pixel patches, grid, resized dimensions and image-token count using the fixed Qwen3.5-4B processor: smart resize, uint8 antialiased bicubic filtering, fused normalization, merge ordering and temporal repetition. Validate dimensions, aspect ratio and exact input length before image allocations.Updated onto main
7f39ac40902c374803992407bb26eeba29c8a588. The six-file diff contains the standalone CPU implementation, its export and the required upstream notices/licenses. No runtime dependency is added. Decoder/HTTP wiring and CUDA vision execution are supplied separately by #64.Test Plan
Current head:
274b082b29d34a73cb0b6a48cb573a540fe732bc. Baseline:7f39ac40902c374803992407bb26eeba29c8a588.Test Result (fresh, 2026-10-05)
The current production tree passed
cargo fmt --all --check,cargo clippy --workspace --locked --all-targets -- -D warnings,cargo test --workspace --locked, andcargo build --workspace --release --lockedon macOS arm64 / Rust 1.98.0. Workspace tests: 82 passed, 0 failed, 10 existing tests ignored (6 CUDA, 2 Laya checkpoint/oracle, 1 Laya packing, 1 Open-Jev tokenization). Benchmark CI checks also passed: 4 smoke requests validated and 8 Python tests passed.Feature fixtures and examples remain outside this core-only diff. A fresh external Cargo runner imports the current production crate by path and uses its locked dependency versions; only the external root package is added to the runner lockfile. The feature harness is materialized from the linked pre-cleanup commit. No native CUDA, full multimodal inference, HTTP image path, model accuracy or performance claim is made by this CPU prerequisite PR.
The external release suite passed all 7 tests, including input rejection, inclusive limits, identity patch order, temporal repetition, normalization, Python ties-to-even geometry and all 14 complete-output golden hashes. The unchanged fixed 24-case held-out corpus was replayed against the actual CPU Hugging Face processor: all 197,468,160 float32 output bytes matched exactly, plus grid, shape, resized dimensions and token counts. Zero tolerance; no resampling or threshold changes.
Reference: processor config revision
851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, config SHA-25627225450ac9c6529872ee1924fcb0962ff5634834f817040f444118116f4e516. Reference versions were checked before replay: Python 3.12, Torch 2.14.0, torchvision 0.29.0, Transformers 5.17.0, NumPy 2.5.3, Pillow 11.3.0. No model weights or GPU resources were used.Harness sources: pre-cleanup commit. Fresh local logs, runner lockfile, frozen protocol, full output hashes and held-out results are archived under
artifacts/all-pr-followups-20261005/cua/pr63/.Self-review
Full current diff reviewed against the pinned main revision; model-specific preprocessing remains separate from forward execution. Required PyTorch, Pillow and Apache notices remain with the adaptation. This CPU preprocessing module can merge independently as a prerequisite for #64. Maintainer review remains separate.
Demo / evidence
Actual CPU validation, not native encoder/HTTP inference. Raw logs and fixture/payload hashes remain local under the artifact path above, outside the core-code diff. No video is needed for these CPU prerequisites. Current-head GitHub Rust, benchmark and docs CI passed; deployment skipped. CI run.