Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
84 changes: 78 additions & 6 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
[workspace]
members = ["src/frontend", "src/runtime", "src/models/clm", "src/models/cua_s1/native", "src/models/qwen3_5/native", "src/models/open_jev/native", "src/models/laya", "src/backends/cuda"]
members = ["src/frontend", "src/runtime", "src/models/clm", "src/models/cua_s1/native", "src/models/qwen3_5/native", "src/models/open_jev/native", "src/models/jev_vl/native", "src/models/laya", "src/backends/cuda"]
resolver = "3"
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -167,6 +167,7 @@ and CLM has a stub-encoder contract recipe:
| Cua-S1 4B 0.2 (`text` adapter) | [Python worker](recipe/cua_s1/text.md); [native worker](recipe/cua_s1/native.md), CUDA, run on sm_89 |
| Cua-S1 4B 0.2 (`multimodal` adapter) | [Python CUDA worker](src/frontend/cua_s1.py); one PNG/JPEG screenshot, `choice`; native screenshot execution remains in progress |
| Open-Jev-27B-v1.1 | [Native Rust/CUDA worker](recipe/open_jev/native.md); eager independent text candidates; [H200 validation](recipe/open_jev/validation.md) |
| autotrust/JEV-27B-VL | [Experimental Rust/CUDA worker](recipe/jev_vl/README.md); single-question text and offline-preencoded image inputs; [bounded H800 validation and limits](recipe/jev_vl/validation.md) |
| CLM-v0.1-8B | [External worker with a CPU stub encoder](recipe/clm/README.md); contract checks only, real Qwen3-8B decisions unverified by this recipe |

[Supported models and hardware](docs/supported-models.md) lists the devices
Expand Down
7 changes: 7 additions & 0 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,6 +59,13 @@ the existing FP32 rounding and bias order. Finishing checks output cardinality
before reconstruction. HTTP validation, error status/body conventions, and real
warmup before readiness remain model-specific and unchanged.

The experimental [JEV-VL worker](../recipe/jev_vl/README.md) also uses the shared
Qwen executor. It admits one official single-decision request at a time, reads
selected LM-head rows, and can retain image-prefix state across requests. Image
features must be prepared offline. Its `{kind, state, question, options}` contract
is distinct from the existing `{model, state, questions}` envelope; frontend
transport alone does not adapt those schemas.

## Layer ownership and implementation language

| Component | Owns | Native target implementation |
Expand Down
4 changes: 3 additions & 1 deletion docs/supported-models.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
# Supported models and hardware

This page covers what runs from `main`. Start with the
This page distinguishes validated serving paths from experimental integrations.
Start with the
[CPU decision walkthrough](getting-started.md). The
[rolling model tracker (#83)](https://github.com/ThinkFlowLab/system1-omni/issues/83)
records available paths separately from proposed integrations and hardware evidence.
Expand All @@ -15,6 +16,7 @@ Models that are being added are also tracked in issues labeled [new model](https
| Cua-S1 4B 0.2, `text` adapter | [Native Rust worker](../recipe/cua_s1/native.md) on the [Qwen3.5 CUDA kernels](../src/backends/cuda/qwen3_5/README.md) | Not supported | Validated on compute capability 8.9 ([#19](https://github.com/ThinkFlowLab/system1-omni/pull/19), [#52](https://github.com/ThinkFlowLab/system1-omni/pull/52)) | Not supported | Compute capability 8.0 or newer, the CUDA toolkit to build, weights merged with `export_text_merged.py` |
| Cua-S1 4B 0.2, `multimodal` adapter | Reference worker on Transformers and PEFT, [`src/frontend/cua_s1.py`](../src/frontend/cua_s1.py); no recipe yet | Not supported | Validated ([#17](https://github.com/ThinkFlowLab/system1-omni/pull/17), [#18](https://github.com/ThinkFlowLab/system1-omni/pull/18)) | Not supported | The state is one PNG or JPEG image; upstream's `weights.lock.json` next to the base weights |
| Open-Jev-27B-v1.1 | [Native Rust/CUDA worker](../recipe/open_jev/native.md) on the shared Qwen3.5/3.8 executor | Not supported | Validated on H200 (sm_90) for the [74 single-candidate workload](../recipe/open_jev/validation.md) | Not supported | Compute capability 8.0 or newer, CUDA toolkit to build, exported merged weights and trained head |
| autotrust/JEV-27B-VL | [Experimental native worker](../recipe/jev_vl/README.md), single-question API | Contract tests only; no CPU inference | H800 frozen-corpus check; [experimental scope and limits](../recipe/jev_vl/validation.md) | Not supported | Pinned merged checkpoint, rebuilt ABI 6 CUDA library; image inputs require offline Transformers preencoding |
| CLM-v0.1-8B | [External `clm-serve` recipe](../recipe/clm/README.md) with a CPU stub embeddings server | **Stub-encoder contract checks only** ([#23](https://github.com/ThinkFlowLab/system1-omni/pull/23)); not real Qwen3-8B decisions | Real encoder unverified by the merged recipe | Unverified | Python, upstream CLM and head checkpoint; a real encoder requires a separate embeddings server |

- **Validated:** covered by the recipe on `main` or by the checks in the linked merged pull request.
Expand Down
2 changes: 2 additions & 0 deletions mkdocs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -67,5 +67,7 @@ nav:
- Cua-S1 native text worker: recipe/cua_s1/native.md
- Open-Jev native text worker: recipe/open_jev/native.md
- Open-Jev H200 validation: recipe/open_jev/validation.md
- JEV-27B-VL experimental worker: recipe/jev_vl/README.md
- JEV-27B-VL validation status: recipe/jev_vl/validation.md
- CLM stub-encoder contract: recipe/clm/README.md
- Contributing: CONTRIBUTING.md
2 changes: 2 additions & 0 deletions recipe/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,8 @@ For a first real decision, follow the [complete CPU walkthrough](../docs/getting
the Rust worker, export the merged weights and start the worker.
- [Open-Jev-27B-v1.1 native text worker](open_jev/native.md): export the merged
text backbone and trained decision head, then serve with Rust and CUDA.
- [JEV-27B-VL experimental worker](jev_vl/README.md): native text and preencoded
image decisions; bounded H800 validation with explicit deployment limits.
- [CLM behind the frontend](clm/README.md): run CLM's own server behind the frontend on
CPU with a stub encoder, and what the response comparison has to allow for.

Expand Down
2 changes: 1 addition & 1 deletion recipe/cua_s1/native.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@ each exact prompt length warms the GEMM plans and captures the forward pass;
later requests replay it with freshly uploaded token ids. At most eight lengths
are cached. Growing the scratch allocation clears the captures before freeing
their buffers. Capture adds first-use latency; leave the variable unset to use
the eager control. Rebuild both the worker and CUDA library together (ABI 4).
the eager control. Rebuild both the worker and CUDA library together (ABI 6).
If capture fails, the worker returns the completed eager result and disables
Graph capture/replay for its remaining lifetime, logging the failure to stderr.

Expand Down
158 changes: 158 additions & 0 deletions recipe/jev_vl/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,158 @@
# JEV-27B-VL experimental native recipe

This recipe serves `autotrust/JEV-27B-VL` System-1 decisions with Rust/CUDA.
Text runs natively; images must first be encoded offline by Transformers.
The [model contract](../../src/models/jev_vl/README.md) explains the single-question
API, verbalizer head and cache ownership. The reviewed candidate passed a
[bounded H800 validation](validation.md#historical-h800-validation);
this remains an experimental integration with the documented coverage limits.

## Prepare a pinned checkpoint

Run all commands from the repository root. Preparation used Linux, Python 3.10,
PyTorch 2.13.0+cu130, Transformers 5.17.0 and safetensors 0.8.0. The
[requirements file](requirements.txt) pins observed direct dependencies; a clean
installation of those pins has not been revalidated. Use an isolated environment:

```sh
python3.10 -m venv .venv-jev-vl
.venv-jev-vl/bin/python -m pip install -r recipe/jev_vl/requirements.txt
.venv-jev-vl/bin/hf download autotrust/JEV-27B-VL \
--revision f34b598d4ef4bcefd337bee8d8e7ddd3b7733ccc \
--local-dir weights/JEV-27B-VL
CUDA_VISIBLE_DEVICES='' .venv-jev-vl/bin/python recipe/jev_vl/export_merged.py \
--model weights/JEV-27B-VL --out weights/jev-vl-merged --max-length 16384
```

The CPU exporter streams shards, merges the backbone LoRA in FP32 before BF16
rounding, and separately exports selected merged LM-head rows as FP32. Keep the
original source checkpoint for image preprocessing. The historical language
export occupied about 48 GiB in addition to the source checkpoint. Peak host
RAM was not measured; the export job requested 64 GiB. Do not assume the whole
pipeline fits on a low-memory workstation from the streaming implementation
alone. The output directory must not exist before export.

## Build and launch

The author measured the worker on one H800 80 GB (`sm_90`). A maintainer also
[reported a bounded L20X replay](https://github.com/ThinkFlowLab/system1-omni/pull/96#issuecomment-6018324526)
on an earlier revision. These runs do not validate later code changes or maximum
context lengths. Use an allocated GPU on
scheduled hosts. Build the CUDA library and Rust workers from the same revision;
the integrated backend uses **ABI 6** and older libraries must be rebuilt.

```sh
src/backends/cuda/qwen3_5/build.sh target/release 90
cargo build --release --locked -p omni-jev-vl-native -p omni-jev
JEV_VL_MODEL=weights/jev-vl-merged \
JEV_VL_CUDA_LIB=$PWD/target/release/libqwen3_5_cuda.so \
JEV_VL_CACHE=0 target/release/omni-jev-vl-native
```

`JEV_VL_HOST` defaults to `127.0.0.1`, `JEV_VL_PORT` to `8001`, and the CUDA
library defaults to the file beside the executable. The worker uses visible
CUDA device 0. `/health` becomes available after model loading and a successful
text warmup; that health check does not validate image assets.

In separate terminals, start the frontend and send the same request directly and
through it:

```sh
OMNI_JEV_BIND=127.0.0.1:8080 OMNI_JEV_BACKEND_URL=http://127.0.0.1:8001 \
target/release/omni-jev
```

```sh
curl --fail-with-body http://127.0.0.1:8001/health
curl --fail-with-body http://127.0.0.1:8080/health
curl --fail-with-body http://127.0.0.1:8001/v1/systemone \
-H 'Content-Type: application/json' --data-binary @recipe/jev_vl/example-request.json
curl --fail-with-body http://127.0.0.1:8080/v1/systemone \
-H 'Content-Type: application/json' --data-binary @recipe/jev_vl/example-request.json
```

Expect a 200 response with two finite `probabilities`, their sum approximately
one, an in-range `choice_index`, and the corresponding string `choice`. Compare
decision fields and usage between both responses; `elapsed_seconds` varies.
This worker takes `kind`, `state`, `question` and `options`, rather than the
`questions`/`answers` envelope in the generic comparison recipe. A fixed expected
choice is not asserted here; use the frozen comparison corpus for numerical
checks. The frontend routes `/health` and `/v1/systemone`; it does not proxy
generation routes such as `/v1/chat/completions`. An unknown route returns the
frontend's own 404 rather than this worker's error envelope.

## Prepare images before serving

The worker does not decode images, fetch URLs or invoke a vision tower on cache
miss. First prepare a JSONL manifest with one object per line containing a
`request` whose `state` includes `{"image":"data:image/png;base64,..."}`. The
preencoder accepts base64 data URIs; a fresh URL or changed image needs a new
asset. It runs on a GPU and should finish before starting the language worker
on the same device:

```sh
.venv-jev-vl/bin/python recipe/jev_vl/preencode.py \
--model weights/JEV-27B-VL --manifest /path/to/requests.jsonl \
--out weights/jev-vl-image-assets
JEV_VL_MODEL=weights/jev-vl-merged JEV_VL_IMGCACHE=weights/jev-vl-image-assets \
JEV_VL_CACHE=1 target/release/omni-jev-vl-native
```

Replace the manifest path with your own prepared workload. The preencoder writes
`sha256(exact_data_uri)/emb.safetensors` and `grid.json`; the worker requires the
exact same URI string. Use a new output directory for a changed checkpoint or
processor. Complete assets are skipped; interrupted writes are regenerated on
the next run. Source identity is not a cryptographic runtime compatibility check. Preencoding,
including image decoding and vision execution, is excluded from the reported
worker timing. This is useful for repeated decisions on a prepared image; it is
not an end-to-end live screenshot service.

## Cache controls

| Variable | Default | Meaning |
| --- | --- | --- |
| `JEV_VL_CACHE` | `1` | Master enable; `0` uses a full language forward and rereads prepared image assets. |
| `JEV_VL_L1`, `JEV_VL_L2`, `JEV_VL_L3` | `1` | Processor records, parsed image assets, and language-prefix state. |
| `JEV_VL_L1_MAX` | `256` | Maximum resident processor records. |
| `JEV_VL_L2_BYTES` | `1073741824` | Parsed image-asset cache budget. |
| `JEV_VL_L3_BYTES` | `2147483648` | Cached language-prefix device-state budget. |

Cache budgets do not include weights, model scratch, in-flight request inputs or
total process memory. `GET /v1/cache/stats` exposes counters and resident cache
accounting. Use those counters to assess L2 hits; `x-jev-cache` does not report
an L2 hit field. `POST /v1/cache/reset` resets counters only; restart the worker for
a cold cache. Request preparation may overlap, while GPU execution stays serial.
Use cache modes as experimental controls within the documented validation scope.

## Local checks

CPU checks do not need weights or CUDA:

```sh
cargo fmt --all --check
cargo clippy --workspace --locked --all-targets -- -D warnings
cargo test --workspace --locked
cargo build --workspace --release --locked
python3 -m unittest discover -s tests/benchmarks -p 'test_jev_vl_*.py' -v
python3 recipe/jev_vl/preencode.py --help
```

The image-prefix tokenizer check requires the exported checkpoint. First
[download and restore the frozen corpus](validation.md#download-the-frozen-corpus).
It uses the corpus's fixed `[1, 60, 60]` grid and synthetic embedding
rows, so it checks token/position splitting rather than the vision encoder.
CUDA kernel tests require an allocated GPU and rebuilt ABI 6 library:

```sh
JEV_VL_EXPORT=$PWD/weights/jev-vl-merged \
JEV_VL_MANIFEST="$jev_vl_evidence/manifest.jsonl" \
cargo test --locked -p omni-jev-vl-native --test jev_vl_prefix -- --ignored
CUA_S1_CUDA_LIB=$PWD/target/release/libqwen3_5_cuda.so \
QWEN3_5_CHECKPOINT=$PWD/weights/jev-vl-merged \
cargo test --release --locked -p omni-qwen3-5-native --test kernels -- \
--ignored --test-threads=1
```

These checks do not replace full-checkpoint parity or direct/frontend HTTP
validation. The [validation page](validation.md) lists the remaining gates and
separates historical measurements from the current revision.
7 changes: 7 additions & 0 deletions recipe/jev_vl/example-request.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
{
"model": "autotrust/JEV-27B-VL",
"kind": "choice",
"state": "The customer received a package with a broken screen and asks for help.",
"question": "What should support do next?",
"options": ["Offer a replacement", "Close the ticket without a reply"]
}
Loading
Loading