Skip to content

cua_s1: preprocess RGB images for native vision on CPU - #63

Merged
hsliuustc0106 merged 4 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-image-preprocess
Oct 6, 2026
Merged

hsliuustc0106 merged 4 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-image-preprocess

Conversation

@Levius-Fubuki

@Levius-Fubuki Levius-Fubuki commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Add image_preprocess::preprocess_rgb8(width, height, rgb) for decoded interleaved RGB8 input. Return float32 pixel patches, grid, resized dimensions and image-token count using the fixed Qwen3.5-4B processor: smart resize, uint8 antialiased bicubic filtering, fused normalization, merge ordering and temporal repetition. Validate dimensions, aspect ratio and exact input length before image allocations.

Updated onto main 7f39ac40902c374803992407bb26eeba29c8a588. The six-file diff contains the standalone CPU implementation, its export and the required upstream notices/licenses. No runtime dependency is added. Decoder/HTTP wiring and CUDA vision execution are supplied separately by #64.

Test Plan

Current head: 274b082b29d34a73cb0b6a48cb573a540fe732bc. Baseline: 7f39ac40902c374803992407bb26eeba29c8a588.

Test Result (fresh, 2026-10-05)

The current production tree passed cargo fmt --all --check, cargo clippy --workspace --locked --all-targets -- -D warnings, cargo test --workspace --locked, and cargo build --workspace --release --locked on macOS arm64 / Rust 1.98.0. Workspace tests: 82 passed, 0 failed, 10 existing tests ignored (6 CUDA, 2 Laya checkpoint/oracle, 1 Laya packing, 1 Open-Jev tokenization). Benchmark CI checks also passed: 4 smoke requests validated and 8 Python tests passed.

Feature fixtures and examples remain outside this core-only diff. A fresh external Cargo runner imports the current production crate by path and uses its locked dependency versions; only the external root package is added to the runner lockfile. The feature harness is materialized from the linked pre-cleanup commit. No native CUDA, full multimodal inference, HTTP image path, model accuracy or performance claim is made by this CPU prerequisite PR.

The external release suite passed all 7 tests, including input rejection, inclusive limits, identity patch order, temporal repetition, normalization, Python ties-to-even geometry and all 14 complete-output golden hashes. The unchanged fixed 24-case held-out corpus was replayed against the actual CPU Hugging Face processor: all 197,468,160 float32 output bytes matched exactly, plus grid, shape, resized dimensions and token counts. Zero tolerance; no resampling or threshold changes.

Reference: processor config revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, config SHA-256 27225450ac9c6529872ee1924fcb0962ff5634834f817040f444118116f4e516. Reference versions were checked before replay: Python 3.12, Torch 2.14.0, torchvision 0.29.0, Transformers 5.17.0, NumPy 2.5.3, Pillow 11.3.0. No model weights or GPU resources were used.

Harness sources: pre-cleanup commit. Fresh local logs, runner lockfile, frozen protocol, full output hashes and held-out results are archived under artifacts/all-pr-followups-20261005/cua/pr63/.

Self-review

Full current diff reviewed against the pinned main revision; model-specific preprocessing remains separate from forward execution. Required PyTorch, Pillow and Apache notices remain with the adaptation. This CPU preprocessing module can merge independently as a prerequisite for #64. Maintainer review remains separate.

Demo / evidence

Actual CPU validation, not native encoder/HTTP inference. Raw logs and fixture/payload hashes remain local under the artifact path above, outside the core-code diff. No video is needed for these CPU prerequisites. Current-head GitHub Rust, benchmark and docs CI passed; deployment skipped. CI run.

@Levius-Fubuki
Levius-Fubuki marked this pull request as ready for review October 5, 2026 15:51
Copilot AI balanced review requested due to automatic review settings October 5, 2026 15:51

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@hsliuustc0106
hsliuustc0106 merged commit 55dcf0a into ThinkFlowLab:main Oct 6, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants