jev_vl: add native inference and prefix caching - #96
linear3735 wants to merge 10 commits into
Conversation
Copy helpers (cs1_copy_dd/copy2d), f32-state GDN continuation (cs1_gdn_prefill_x) and windowed gated attention (cs1_attention_gated_prefix) so a forward can resume from a precomputed prefix state bitwise-equivalent to a one-shot pass (except GEMM M-shape). Library ABI stays at version 4; existing workers are unaffected. Evidence and conditions: ThinkFlowLab/system1-omni run /aifs4su/shiheming/jev/runs/20261006-jev-vl-native-r2 (H800, single-pass versus two-stage bitwise checks in tests/qwen3_5/kernels.rs). Co-Authored-By: Claude Code <noreply@anthropic.com>
Add inputs.rs (multimodal prompt assembly: placeholder rows overlaid with precomputed bf16 image embeddings plus custom 3D-position rotary tables via HF get_rope_index semantics) and PrefixState capture/continue/run_window in model.rs (f32 state, window-aligned to the 64-token GDN chunking). Existing text-only callers keep the same behavior; kernel-side occurrences are covered by tests/qwen3_5/kernels.rs. Co-Authored-By: Claude Code <noreply@anthropic.com>
Serve autotrust/JEV-27B-VL System-1 decisions through /v1/systemone with a new omni-jev-vl-native crate: streaming LoRA merge export (backbone and lm_head merged separately, 256 verbalizer label rows), official (logprob(label)+bias)/per-kind-temperature readout semantics, and three Rust-side caches (L1 structural prefix, L2 vision-encoder assets by media hash, L3 KV prefix captured per structure key). Open-Jev and Cua-S1 sources are untouched. Measured on one H800 GPU, warm, two passes against the official vLLM serve path (R1 frozen reference): text manifest 36/36 and image set 12/12 decisions identical (48/48, max |dp| <= 0.0215); single-request p50 28.2ms vs 94.1ms; same-image multi-question cache hit vs miss p50 34.5ms vs 104.2ms (3.02x), with L3 prefix reuse the dominant term. Full protocol, raw samples and conditions: /aifs4su/shiheming/jev/runs/20261006-jev-vl-native-r2/. Co-Authored-By: Claude Code <noreply@anthropic.com>
hsliuustc0106
left a comment
There was a problem hiding this comment.
Reviewed 522f2256876a62ddd70873ee9e775055416aea1b against merge base 47eff9cdeda01e4847a4fb9634a43f2cab6a233f. Found two P2 validation defects, detailed in the inline comments and confirmed with local CPU probes.
Local validation:
cargo fmt --all --check: passed.cargo clippy --workspace --locked --all-targets -- -D warnings: passed.cargo test --workspace --locked: 121 passed, 19 ignored.cargo build --workspace --release --locked: passed.python -m unittest discover -s tests/benchmarks -p 'test_*.py' -v: 14 passed.mkdocs build --strict: passed with the pinned documentation dependencies.- Qwen CUDA kernel tests: 11 passed on one reserved NVIDIA L20X (
sm_89), using a freshly built ABI 6 CUDA library. This includes bitwise windowed-attention and two-stage GDN continuation checks. The checkpoint-dependent prefix-ownership test was excluded. manifest_suffix_split_matches_full_expand --ignored: passed using the pinned released tokenizer, its generated label table, and the exact restored frozen manifest; no checkpoint weights were required.
Seven checkpoint/oracle tests remain skipped after the additional kernel and suffix tests. Full JEV-VL checkpoint parity and the H800 latency comparison were not rerun because the prepared merged weights and image assets were unavailable locally.
The GPU test runner printed a successful result for all 11 tests. Its PRoot/scheduler wrapper subsequently hung while waiting for the supervisor, so I cancelled only that task-owned run after test completion (wrapper exit 137). The reservation was released and no task GPU processes remained. The reviewed source was unchanged, and the PR head was rechecked before posting.
| return Err(Reject::bad_request("question must be a string".to_owned())); | ||
| } | ||
| }; | ||
| if let Some(thinking) = request.get("thinking").and_then(Value::as_str) { |
There was a problem hiding this comment.
[P2] Reject non-string decision controls
thinking:true, thinking:null, and strategy:42 all pass compile() because and_then(Value::as_str) converts a present non-string value into None. The request then follows the default System1 path rather than rejecting the invalid control. The pinned reference's DecideRequest rejects each of these with literal_error; I checked that reference class and reproduced acceptance in the native compiler locally.
Validate the type of each present control before applying defaults, including the analogous strategy check below.
| let tail_ids = self.tokenize(&tail)?; | ||
| let suffix_pads = pads_end - p; | ||
| let mut ids = Vec::with_capacity(suffix_pads + 1 + tail_ids.len()); | ||
| ids.extend(std::iter::repeat_n(self.image_pad, suffix_pads)); | ||
| ids.push(self.vision_end); | ||
| ids.extend_from_slice(&tail_ids); |
There was a problem hiding this comment.
[P2] Validate image placeholders on cache hits
Warm L1 with a valid single-image request, then reuse the same kind/state with a question containing <|image_pad|>. This hit path appends the extra placeholder from tail_ids without including an embedding/index for it, and returns a prepared plan. MultimodalInput::validate() then fails in the executor, which the HTTP handler maps to 500. With caching disabled, images::expand() catches the same placeholder mismatch during preparation and returns 400.
I reproduced the preparation/validation difference on CPU using the production processor and input validator, the real pinned tokenizer, synthetic image/head fixtures, and L3 disabled. Reject unexpected image placeholders in the fresh suffix before dispatch so cache state does not change validation behavior.
|
Follow-up local performance validation for
The corresponding pass speedups were 3.33× and 3.29×; the combined median ratio was 3.312×. Each mode used two measured passes of the same 12 questions about one prepared synthetic image, one serial client, and 12 excluded warmup requests before each pass. One server per configuration served both passes. Controls: same physical GPU 4, release worker binary, ABI 6 CUDA library, BF16 merged checkpoint and prepared image asset; CPU affinity Validation:
Preparation used This is a bounded local L20X cache comparison, not a new H800 or vLLM baseline, a concurrency benchmark, or a cold-start/online-vision result. Provenance and raw measured latenciesManifest SHA-256: Worker SHA-256: CUDA library SHA-256: Milliseconds, in frozen request order: |
Signed-off-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Purpose
Add a Rust/CUDA worker for
autotrust/JEV-27B-VLdecisions:choice,noulandscore. It reuses prepared image embeddings and language-prefix state acrossquestions about the same image.
The worker uses the shared Qwen backend and serial runtime. It requires CUDA
ABI 6; rebuild the library and dependent workers together. Vision encoding runs
offline. The API accepts
{kind,state,question,options}.Test Plan
System1-Omni Version / Commit: merge base
9986574, head37525bc.The full diff is 44 files, +5,649/-46 lines. Authored code accounts for 4,928
changed lines: 2,770 source, 1,662 tests and 496 preparation/validation scripts.
Documentation, licenses, configuration, lockfile and the request example are
excluded from the code count.
Review in this order:
recipe/jev_vl/src/models/jev_vl/src/models/qwen3_5/native/,src/backends/cuda/qwen3_5/The shared backend could be split into a dependency PR. It stays here because
its new entry points serve this worker's prepared-image and prefix-cache path.
Keeping them together provides one runnable feature; no separate optimization
or unrelated refactor is included. Run output is archived outside the source
tree. The only added JSON file is the seven-line request example.
Test Result
Local checks passed:
Rust: 125 passed, 20 ignored. Python: 23 passed, using the benchmark
requirements. Documentation used
docs/requirements.txt. The new image-writetests reproduced failures before the fix and passed afterward.
GitHub CI for
37525bcpassed: Rust and benchmarksand documentation build.
This revision merges upstream #102, preserving its packed text path and the
separate JEV prefix path.
No new GPU measurements were run for the cleanup, input-boundary fixes or
upstream merge. The combined ABI 6 packing/Graph and JEV prefix paths still
need checkpoint regression; #102's historical ABI 5 results do not cover this merge.
Demo / evidence
Deployment recipe
and validation/reproduction
include setup commands, the fixed corpus and archived raw results.
Historical H800 checks matched 48/48 frozen decisions in each cache mode within
the fixed 0.025 probability tolerance. Two passes of 12 prepared-image questions
per mode measured warm worker-direct HTTP p50 at 102.88 ms off and 37.04 ms on
(2.78×), with 12 excluded warmups per pass. Timing includes JSON decoding;
it excludes offline vision, startup and frontend forwarding. This is a cache
comparison at concurrency 1, not a current-head measurement or a vLLM speedup.
A maintainer separately reported a bounded L20X replay on
522f225.Both records identify their measured revisions. Clean installation, current-head
GPU parity, broad model quality, peak memory, multi-client performance and full
Cua-S1/Open-Jev checkpoint regression remain unverified.
Self-review