Skip to content

jev_vl: add native inference and prefix caching - #96

Open
linear3735 wants to merge 10 commits into
ThinkFlowLab:mainfrom
linear3735:codex/jev-vl-native-mm-cache
Open

linear3735 wants to merge 10 commits into
ThinkFlowLab:mainfrom
linear3735:codex/jev-vl-native-mm-cache

Conversation

@linear3735

@linear3735 linear3735 commented Oct 6, 2026 •

Copy link
Copy Markdown

Purpose

Add a Rust/CUDA worker for autotrust/JEV-27B-VL decisions: choice, noul and
score. It reuses prepared image embeddings and language-prefix state across
questions about the same image.

The worker uses the shared Qwen backend and serial runtime. It requires CUDA
ABI 6; rebuild the library and dependent workers together. Vision encoding runs
offline. The API accepts {kind,state,question,options}.

Test Plan

System1-Omni Version / Commit: merge base 9986574, head 37525bc.

The full diff is 44 files, +5,649/-46 lines. Authored code accounts for 4,928
changed lines
: 2,770 source, 1,662 tests and 496 preparation/validation scripts.
Documentation, licenses, configuration, lockfile and the request example are
excluded from the code count.

Review in this order:

Area Responsibility Validation at this head
recipe/jev_vl/ Merge the trained weights, prepare image assets, replay requests 15 JEV Python tests pass, including interrupted-write recovery. Exporter syntax checked; real export and vision execution not rerun.
src/models/jev_vl/ Parse requests, prepare model inputs, read decision scores, own caches 35 CPU tests pass; one checkpoint-dependent tokenizer/prefix check is ignored.
src/models/qwen3_5/native/, src/backends/cuda/qwen3_5/ Image-row overwrite, 3D rotary and prefix capture/continue 5 Qwen CPU tests pass; 12 CUDA/checkpoint tests compile but are ignored. New prefix-length cases require a GPU checkpoint run.
Shared worker integration and docs Reuse serial dispatch, document ABI and deployment Workspace checks, frontend tests, benchmark smoke and strict MkDocs pass. Current-head live worker/frontend replay remains unverified.

The shared backend could be split into a dependency PR. It stays here because
its new entry points serve this worker's prepared-image and prefix-cache path.
Keeping them together provides one runnable feature; no separate optimization
or unrelated refactor is included. Run output is archived outside the source
tree. The only added JSON file is the seven-line request example.

Test Result

Local checks passed:

cargo fmt --all --check
cargo clippy --workspace --locked --all-targets -- -D warnings
cargo test --workspace --locked
cargo build --workspace --release --locked
python -m unittest discover -s tests/benchmarks -p 'test_*.py' -v
python benchmarks/bench.py validate benchmarks/smoke.jsonl
python -S recipe/jev_vl/preencode.py --help
python -m mkdocs build --strict

Rust: 125 passed, 20 ignored. Python: 23 passed, using the benchmark
requirements. Documentation used docs/requirements.txt. The new image-write
tests reproduced failures before the fix and passed afterward.

GitHub CI for 37525bc passed: Rust and benchmarks
and documentation build.
This revision merges upstream #102, preserving its packed text path and the
separate JEV prefix path.

No new GPU measurements were run for the cleanup, input-boundary fixes or
upstream merge. The combined ABI 6 packing/Graph and JEV prefix paths still
need checkpoint regression; #102's historical ABI 5 results do not cover this merge.

Demo / evidence

Deployment recipe
and validation/reproduction
include setup commands, the fixed corpus and archived raw results.

Historical H800 checks matched 48/48 frozen decisions in each cache mode within
the fixed 0.025 probability tolerance. Two passes of 12 prepared-image questions
per mode measured warm worker-direct HTTP p50 at 102.88 ms off and 37.04 ms on
(2.78×), with 12 excluded warmups per pass. Timing includes JSON decoding;
it excludes offline vision, startup and frontend forwarding. This is a cache
comparison at concurrency 1, not a current-head measurement or a vLLM speedup.

A maintainer separately reported a bounded L20X replay on 522f225.
Both records identify their measured revisions. Clean installation, current-head
GPU parity, broad model quality, peak memory, multi-client performance and full
Cua-S1/Open-Jev checkpoint regression remain unverified.

Self-review

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

linear3735 and others added 4 commits October 6, 2026 15:38
Copy helpers (cs1_copy_dd/copy2d), f32-state GDN continuation
(cs1_gdn_prefill_x) and windowed gated attention
(cs1_attention_gated_prefix) so a forward can resume from a
precomputed prefix state bitwise-equivalent to a one-shot pass
(except GEMM M-shape). Library ABI stays at version 4; existing
workers are unaffected.

Evidence and conditions: ThinkFlowLab/system1-omni run
/aifs4su/shiheming/jev/runs/20261006-jev-vl-native-r2 (H800,
single-pass versus two-stage bitwise checks in
tests/qwen3_5/kernels.rs).

Co-Authored-By: Claude Code <noreply@anthropic.com>
Add inputs.rs (multimodal prompt assembly: placeholder rows
overlaid with precomputed bf16 image embeddings plus custom
3D-position rotary tables via HF get_rope_index semantics) and
PrefixState capture/continue/run_window in model.rs (f32 state,
window-aligned to the 64-token GDN chunking). Existing text-only
callers keep the same behavior; kernel-side occurrences are
covered by tests/qwen3_5/kernels.rs.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Serve autotrust/JEV-27B-VL System-1 decisions through /v1/systemone
with a new omni-jev-vl-native crate: streaming LoRA merge export
(backbone and lm_head merged separately, 256 verbalizer label rows),
official (logprob(label)+bias)/per-kind-temperature readout
semantics, and three Rust-side caches (L1 structural prefix,
L2 vision-encoder assets by media hash, L3 KV prefix captured per
structure key). Open-Jev and Cua-S1 sources are untouched.

Measured on one H800 GPU, warm, two passes against the official
vLLM serve path (R1 frozen reference): text manifest 36/36 and
image set 12/12 decisions identical (48/48, max |dp| <= 0.0215);
single-request p50 28.2ms vs 94.1ms; same-image multi-question
cache hit vs miss p50 34.5ms vs 104.2ms (3.02x), with L3 prefix
reuse the dominant term. Full protocol, raw samples and conditions:
/aifs4su/shiheming/jev/runs/20261006-jev-vl-native-r2/.

Co-Authored-By: Claude Code <noreply@anthropic.com>

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed 522f2256876a62ddd70873ee9e775055416aea1b against merge base 47eff9cdeda01e4847a4fb9634a43f2cab6a233f. Found two P2 validation defects, detailed in the inline comments and confirmed with local CPU probes.

Local validation:

  • cargo fmt --all --check: passed.
  • cargo clippy --workspace --locked --all-targets -- -D warnings: passed.
  • cargo test --workspace --locked: 121 passed, 19 ignored.
  • cargo build --workspace --release --locked: passed.
  • python -m unittest discover -s tests/benchmarks -p 'test_*.py' -v: 14 passed.
  • mkdocs build --strict: passed with the pinned documentation dependencies.
  • Qwen CUDA kernel tests: 11 passed on one reserved NVIDIA L20X (sm_89), using a freshly built ABI 6 CUDA library. This includes bitwise windowed-attention and two-stage GDN continuation checks. The checkpoint-dependent prefix-ownership test was excluded.
  • manifest_suffix_split_matches_full_expand --ignored: passed using the pinned released tokenizer, its generated label table, and the exact restored frozen manifest; no checkpoint weights were required.

Seven checkpoint/oracle tests remain skipped after the additional kernel and suffix tests. Full JEV-VL checkpoint parity and the H800 latency comparison were not rerun because the prepared merged weights and image assets were unavailable locally.

The GPU test runner printed a successful result for all 11 tests. Its PRoot/scheduler wrapper subsequently hung while waiting for the supervisor, so I cancelled only that task-owned run after test completion (wrapper exit 137). The reservation was released and no task GPU processes remained. The reviewed source was unchanged, and the PR head was rechecked before posting.

return Err(Reject::bad_request("question must be a string".to_owned()));
}
};
if let Some(thinking) = request.get("thinking").and_then(Value::as_str) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Reject non-string decision controls

thinking:true, thinking:null, and strategy:42 all pass compile() because and_then(Value::as_str) converts a present non-string value into None. The request then follows the default System1 path rather than rejecting the invalid control. The pinned reference's DecideRequest rejects each of these with literal_error; I checked that reference class and reproduced acceptance in the native compiler locally.

Validate the type of each present control before applying defaults, including the analogous strategy check below.

Comment on lines +190 to +195
let tail_ids = self.tokenize(&tail)?;
let suffix_pads = pads_end - p;
let mut ids = Vec::with_capacity(suffix_pads + 1 + tail_ids.len());
ids.extend(std::iter::repeat_n(self.image_pad, suffix_pads));
ids.push(self.vision_end);
ids.extend_from_slice(&tail_ids);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Validate image placeholders on cache hits

Warm L1 with a valid single-image request, then reuse the same kind/state with a question containing <|image_pad|>. This hit path appends the extra placeholder from tail_ids without including an embedding/index for it, and returns a prepared plan. MultimodalInput::validate() then fails in the executor, which the HTTP handler maps to 500. With caching disabled, images::expand() catches the same placeholder mismatch during preparation and returns 400.

I reproduced the preparation/validation difference on CPU using the production processor and input validator, the real pinned tokenizer, synthetic image/head fixtures, and L3 disabled. Reject unexpected image placeholders in the fresh suffix before dispatch so cache state does not change validation behavior.

Copy link
Copy Markdown
Contributor

Follow-up local performance validation for 522f2256876a62ddd70873ee9e775055416aea1b: 3.31× faster warm worker-direct HTTP p50 with caching enabled on one NVIDIA L20X (sm_89).

Mode Pass 1 p50 Pass 2 p50 Combined p50
JEV_VL_CACHE=0 104.756 ms 103.884 ms 104.406 ms
JEV_VL_CACHE=1 31.423 ms 31.534 ms 31.523 ms

The corresponding pass speedups were 3.33× and 3.29×; the combined median ratio was 3.312×. Each mode used two measured passes of the same 12 questions about one prepared synthetic image, one serial client, and 12 excluded warmup requests before each pass. One server per configuration served both passes.

Controls: same physical GPU 4, release worker binary, ABI 6 CUDA library, BF16 merged checkpoint and prepared image asset; CPU affinity 56-63 and memory binding to NUMA node 1; max length 16,384; CUA_S1_GRAPH=0; L1/L2/L3 enabled, varying only JEV_VL_CACHE. Both workers used identical PRoot isolation. Timing includes worker-direct loopback HTTP and response JSON decoding, and excludes download/copy/export, offline vision encoding, startup, feasibility and warmup. Shared caches were not dropped.

Validation:

  • All 48 measured requests matched the frozen decisions and passed the fixed 0.025 probability gate, with no failures. Maximum absolute probability differences were 0.021545 off / 0.014215 on.
  • The separate full 48-request frozen-corpus feasibility run also passed. Readiness followed the worker's real warmup; the first inference after readiness was validated in each configuration.
  • Cache-on counters across warmup + measurement showed 3 prefix populations, 45 L3 hits and 0 full-forward fallbacks. Cache-off retained no prefix states.
  • The GPU run exited 0 and its reservation was released; no task GPU processes remained.

Preparation used autotrust/JEV-27B-VL@f34b598d4ef4bcefd337bee8d8e7ddd3b7733ccc and the recipe's pinned PyTorch 2.13.0+cu130, Transformers 5.17.0, tokenizers 0.23.2, safetensors 0.8.0, HF Hub 1.33.0 and Pillow 12.3.0. All 18 reused base shards were content-hash verified against the pinned release; missing adapter/configuration files were downloaded. An initial feasibility attempt using older preparation dependencies exceeded the probability tolerance on one image question (0.033582), although all decisions matched. That failed attempt was retained and excluded; rebuilding the image preparation with the pinned environment resolved it. The language-export file hashes were unchanged.

This is a bounded local L20X cache comparison, not a new H800 or vLLM baseline, a concurrency benchmark, or a cold-start/online-vision result.

Provenance and raw measured latencies

Manifest SHA-256: f72d1beaaaf53933d0a6edda26b635f46931990d8cdb56ca5c6a7ca94d2eb0ee

Worker SHA-256: b7fac1150c7f9f6716ffaf8251455e6f4c5257ca613aac9bb2856e1eefec35af

CUDA library SHA-256: 85fa5e8de2ab58ec12af268251e51a874ba667811fe800daf2e1034d98d599b8

Milliseconds, in frozen request order: img-noul-01..04, img-score-01..04, img-choice-01..04. Warmup and feasibility samples are excluded.

Off pass 1: [103.282530, 102.051489, 104.181448, 104.881464, 104.928924, 104.631157, 102.741496, 103.705611, 108.858632, 109.069110, 109.024969, 106.993222]
Off pass 2: [103.646757, 103.611528, 102.503474, 102.616041, 101.700633, 104.121456, 104.954227, 103.527007, 107.594832, 107.588288, 109.956579, 108.425877]
On pass 1:  [31.559023, 31.452288, 31.276129, 31.092942, 31.301343, 31.263876, 31.169955, 31.392904, 32.132128, 32.762332, 33.325966, 32.369485]
On pass 2:  [31.178059, 31.301714, 31.524602, 31.274517, 31.542856, 31.521992, 31.690842, 31.286560, 32.185652, 32.490521, 33.096546, 31.847882]

Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants