Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ Read `src/models/open_jev/README.md`, `src/models/open_jev/native/src/contract.r

- The reference compiler/readout is pinned to Open-Jev `3308a15ccd7eea1df7a37d6ddc39b023b801ba16`; backbone and checkpoint identities are checked in the exported manifest. Cua-S1's prompt, option limit, JSON ordering, and confidence formula do not apply to this model.
- Each candidate has an independent prompt and scalar head score. Preserve candidate order and sorted structured-input rendering. Normalize across each complete question using the saved temperature; `noul` uses logits `[0, score]`. Candidate regrouping must preserve question identity and token usage.
- The worker validates and tokenizes every candidate before inference, warms up before binding, and uses the shared Qwen executor/CUDA ABI. Current scoring is independent single-prompt execution with a CPU head. Use `recipe/open_jev/validation.md` for the scope and limitations of full-checkpoint comparisons.
- The worker validates and tokenizes every candidate before inference, warms up before binding, and uses the shared Qwen executor/CUDA ABI. Current scoring is independent single-prompt execution with a CPU head. Use `recipe/open_jev/validation.md` (27B on H200) and `recipe/open_jev/validation-9b.md` (9B) for the scope and limitations of full-checkpoint comparisons.

## CUDA, graph and compatibility evidence

Expand Down
15 changes: 8 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,8 +43,8 @@ scheduling from model execution; a native Metal backend is planned, while LAYA
already has a Python worker for Apple GPUs through PyTorch MPS.

The Rust frontend forwards requests to a separately running model worker. The
Cua-S1 4B 0.2 `text` adapter and Open-Jev-27B-v1.1 have native workers using
shared CUDA kernels in this repository.
Cua-S1 4B 0.2 `text` adapter, Open-Jev-27B-v1.1 and Open-Jev-9B have native
workers using shared CUDA kernels in this repository.

## News

Expand All @@ -64,9 +64,9 @@ shared CUDA kernels in this repository.
learned heads, device state, and kernel selection. Native workers have separate
processing and executor modules, with shared FIFO admission and blocking
dispatch in the [native runtime](src/runtime/README.md).
- **Native CUDA workers.** The Cua-S1 4B 0.2 `text` adapter and
Open-Jev-27B-v1.1 run as native workers with shared CUDA kernels. Cua-S1
also has a Python worker that serves as the correctness reference.
- **Native CUDA workers.** The Cua-S1 4B 0.2 `text` adapter,
Open-Jev-27B-v1.1 and Open-Jev-9B run as native workers with shared CUDA
kernels. Cua-S1 also has a Python worker that serves as the correctness reference.
- **LAYA text serving.** LAYA runs as an external CPU Python worker, the
in-repository Python MPS/CPU worker, or a native Rust/CUDA worker on Hopper.
- **CUDA backend and planned Metal backend.** High-performance GPU operations
Expand Down Expand Up @@ -158,15 +158,16 @@ The [frontend documentation](src/frontend/README.md) describes transport and con
LAYA text serving uses the upstream CPU worker, the in-repository Python MPS/CPU
worker, or a native Rust/CUDA worker on Hopper. The Cua-S1 4B 0.2 `text` adapter
runs as a Python worker or as a native worker on CUDA. Open-Jev-27B-v1.1
runs as a native Rust/CUDA worker. Cua-S1 also has a Python screenshot worker,
and CLM has a stub-encoder contract recipe:
and Open-Jev-9B run on the same native Rust/CUDA worker. Cua-S1 also has a
Python screenshot worker, and CLM has a stub-encoder contract recipe:

| Model | Status |
| --- | --- |
| LAYA | [External worker](recipe/laya/README.md); [Python worker on Apple Silicon (MPS) and CPU](recipe/laya/apple-silicon.md); [CPU checkpoint reader](src/models/laya/README.md); [native Rust/CUDA worker on Hopper](recipe/laya/native/README.md) |
| Cua-S1 4B 0.2 (`text` adapter) | [Python worker](recipe/cua_s1/text.md); [native worker](recipe/cua_s1/native.md), CUDA, run on sm_89 |
| Cua-S1 4B 0.2 (`multimodal` adapter) | [Python CUDA worker](src/frontend/cua_s1.py); one PNG/JPEG screenshot, `choice`; native screenshot execution remains in progress |
| Open-Jev-27B-v1.1 | [Native Rust/CUDA worker](recipe/open_jev/native.md); eager independent text candidates; [H200 validation](recipe/open_jev/validation.md) |
| Open-Jev-9B | The same [native Rust/CUDA worker](recipe/open_jev/native.md); eager independent text candidates; [reference comparison on sm_89](recipe/open_jev/validation-9b.md) |
| CLM-v0.1-8B | [External worker with a CPU stub encoder](recipe/clm/README.md); contract checks only, real Qwen3-8B decisions unverified by this recipe |

[Supported models and hardware](docs/supported-models.md) lists the devices
Expand Down
12 changes: 7 additions & 5 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,9 +9,10 @@ design. Concrete input/output types follow each executor's supported layout.
The [Rust frontend](../src/frontend/README.md) currently forwards HTTP requests
to separately running workers. Cua-S1 and Open-Jev have native Rust/CUDA workers
that share the [Qwen3.5/3.8 executor](../src/models/qwen3_5/native/), which accepts
single prompts and bounded packed prefill. Cua-S1 uses single-prompt calls;
Open-Jev packs candidates within one request for input and gate/up GEMMs while
preserving per-sequence mixers and output/down GEMM shapes.
single prompts and bounded packed prefill. Cua-S1 and Open-Jev-9B use
single-prompt calls; Open-Jev-27B-v1.1 packs candidates within one request for
input and gate/up GEMMs while preserving per-sequence mixers and output/down GEMM
shapes.
[Laya's native worker](../src/models/laya/README.md)
uses a separate Hopper CUDA backend for one complete padded request. All three
coordinate independent processors and executors through
Expand Down Expand Up @@ -40,8 +41,9 @@ and candidate identity, usage, and response metadata outside the executor.
| Open-Jev | Token-ID vectors grouped by question, then independent candidate, in request order. | One FP32 learned scalar per candidate in the same grouping. | Add the `noul` false logit of zero, calibrate across each complete question, and restore typed answers, usage, and metadata. |
| Laya | One padded request: token IDs, true lengths, question types and ordered option markers; at most 16 questions, 512 tokens per row and 2048 markers. | Per-question FP32 option logits and two action logits copied back after GPU heads. | Calibrate and decode ordered `choice`, `score` and `noul` answers, usage and metadata. |

Cua-S1 input collections are serial work. Open-Jev's model-specific batch adapter
packs up to 16 independent candidates and 4096 tokens per group; longer prompts
Cua-S1 and Open-Jev-9B input collections are serial work. For Open-Jev-27B-v1.1,
Open-Jev's model-specific batch adapter packs up to 16 independent candidates and
4096 tokens per group; longer prompts
execute alone. It restores question/candidate grouping before normalization.
Laya batches questions within one request. Shared runtime
admission precedes blocking dispatch: Cua-S1 admits one question forward at a
Expand Down
1 change: 1 addition & 0 deletions docs/supported-models.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,7 @@ Models that are being added are also tracked in issues labeled [new model](https
| Cua-S1 4B 0.2, `text` adapter | [Native Rust worker](../recipe/cua_s1/native.md) on the [Qwen3.5 CUDA kernels](../src/backends/cuda/qwen3_5/README.md) | Not supported | Validated on compute capability 8.9 ([#19](https://github.com/ThinkFlowLab/system1-omni/pull/19), [#52](https://github.com/ThinkFlowLab/system1-omni/pull/52)) | Not supported | Compute capability 8.0 or newer, the CUDA toolkit to build, weights merged with `export_text_merged.py` |
| Cua-S1 4B 0.2, `multimodal` adapter | Reference worker on Transformers and PEFT, [`src/frontend/cua_s1.py`](../src/frontend/cua_s1.py); no recipe yet | Not supported | Validated ([#17](https://github.com/ThinkFlowLab/system1-omni/pull/17), [#18](https://github.com/ThinkFlowLab/system1-omni/pull/18)) | Not supported | The state is one PNG or JPEG image; upstream's `weights.lock.json` next to the base weights |
| Open-Jev-27B-v1.1 | [Native Rust/CUDA worker](../recipe/open_jev/native.md) on the shared Qwen3.5/3.8 executor | Not supported | Validated on H200 (sm_90) for the [74 single-candidate workload](../recipe/open_jev/validation.md) | Not supported | Compute capability 8.0 or newer, CUDA toolkit to build, exported merged weights and trained head |
| Open-Jev-9B | The same [native Rust/CUDA worker](../recipe/open_jev/native.md) | Not supported | Validated on compute capability 8.9 against the reference for [253 requests](../recipe/open_jev/validation-9b.md) | Not supported | Compute capability 8.0 or newer, CUDA toolkit to build, exported merged weights and trained head |
| CLM-v0.1-8B | [External `clm-serve` recipe](../recipe/clm/README.md) with a CPU stub embeddings server | **Stub-encoder contract checks only** ([#23](https://github.com/ThinkFlowLab/system1-omni/pull/23)); not real Qwen3-8B decisions | Real encoder unverified by the merged recipe | Unverified | Python, upstream CLM and head checkpoint; a real encoder requires a separate embeddings server |

- **Validated:** covered by the recipe on `main` or by the checks in the linked merged pull request.
Expand Down
3 changes: 2 additions & 1 deletion mkdocs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -66,6 +66,7 @@ nav:
- Cua-S1 text worker: recipe/cua_s1/text.md
- Cua-S1 native text worker: recipe/cua_s1/native.md
- Open-Jev native text worker: recipe/open_jev/native.md
- Open-Jev H200 validation: recipe/open_jev/validation.md
- Open-Jev-27B H200 validation: recipe/open_jev/validation.md
- Open-Jev-9B validation: recipe/open_jev/validation-9b.md
- CLM stub-encoder contract: recipe/clm/README.md
- Contributing: CONTRIBUTING.md
6 changes: 4 additions & 2 deletions recipe/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,8 +13,10 @@ For a first real decision, follow the [complete CPU walkthrough](../docs/getting
the worker and connect the Rust frontend.
- [Cua-S1 4B 0.2 native text worker](cua_s1/native.md): build the CUDA library and
the Rust worker, export the merged weights and start the worker.
- [Open-Jev-27B-v1.1 native text worker](open_jev/native.md): export the merged
text backbone and trained decision head, then serve with Rust and CUDA.
- [Open-Jev native text worker](open_jev/native.md): export the merged text
backbone and trained decision head of Open-Jev-27B-v1.1 or Open-Jev-9B, then
serve with Rust and CUDA. [Open-Jev-9B validation](open_jev/validation-9b.md)
compares the 9B worker with the reference.
- [CLM behind the frontend](clm/README.md): run CLM's own server behind the frontend on
CPU with a stub encoder, and what the response comparison has to allow for.

Expand Down
34 changes: 24 additions & 10 deletions recipe/open_jev/export_merged.py
Original file line number Diff line number Diff line change
@@ -1,7 +1,8 @@
"""Export the pinned Open-Jev-27B-v1.1 text backbone and scalar head on CPU.
"""Export a pinned Open-Jev text backbone and scalar head on CPU.

Use the reference environment documented in native.md. This preparation step
needs about 110 GB of host RAM and 52 GB of output storage, without a GPU.
runs without a GPU. Open-Jev-27B-v1.1 needs about 110 GB of host RAM and 52 GB
of output storage; Open-Jev-9B needs about 20 GB of RAM and 16 GB of storage.
"""

import argparse
Expand All @@ -12,8 +13,19 @@
from peft import PeftModel
from transformers import AutoModelForImageTextToText, AutoTokenizer

BASE_REVISION = "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0"
CHECKPOINT_REVISION = "28cf73067d5b337860bbef3c85b8b82ba8730956"
# Base model -> (base revision, checkpoint revision, scalar head width).
CHECKPOINTS = {
"Qwen/Qwen3.8-27B": (
"1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0",
"28cf73067d5b337860bbef3c85b8b82ba8730956",
5120,
),
"Qwen/Qwen3.5-9B": (
"c202236235762e1c871ad0ccb60c8ee5ba337b9a",
"47e966881e489511c0c7f5633a9e1960a676a551",
4096,
),
}


def main():
Expand All @@ -24,8 +36,10 @@ def main():
parser.add_argument("--max-length", type=int, default=4096)
args = parser.parse_args()
config = json.loads((args.checkpoint / "model.json").read_text())
if config["model_id"] != "Qwen/Qwen3.8-27B" or config["revision"] != BASE_REVISION:
raise ValueError("expected Open-Jev-27B-v1.1's pinned base")
pins = CHECKPOINTS.get(config["model_id"])
if pins is None or config["revision"] != pins[0]:
raise ValueError("expected Open-Jev-27B-v1.1's or Open-Jev-9B's pinned base")
base_revision, checkpoint_revision, width = pins
if args.out.exists():
raise ValueError("output already exists; choose a new export directory")
if not 1 <= args.max_length <= 16384:
Expand All @@ -41,8 +55,8 @@ def main():
raise ValueError("expected a single-user text chat template")
prefix, suffix = chat.split(marker)
head = torch.load(args.checkpoint / "head.pt", map_location="cpu", weights_only=True)
if head["weight"].shape != (1, 5120) or head["bias"].shape != (1,):
raise ValueError("expected a 5120-wide trained scalar head")
if head["weight"].shape != (1, width) or head["bias"].shape != (1,):
raise ValueError(f"expected a {width}-wide trained scalar head")
if not all(torch.isfinite(v).all() for v in head.values()):
raise ValueError("non-finite scalar head")
full = AutoModelForImageTextToText.from_pretrained(
Expand All @@ -58,8 +72,8 @@ def main():
# Written last: the native worker refuses incomplete exports or plain base weights.
(args.out / "open_jev_export.json").write_text(json.dumps({
"format": "open-jev-text-merged/1",
"model_id": config["model_id"], "base_revision": BASE_REVISION,
"checkpoint_revision": CHECKPOINT_REVISION, "temperature": temperature,
"model_id": config["model_id"], "base_revision": base_revision,
"checkpoint_revision": checkpoint_revision, "temperature": temperature,
"max_length": args.max_length, "chat_prefix": prefix, "chat_suffix": suffix,
"head_weight": head["weight"].float().reshape(-1).tolist(),
"head_bias": head["bias"].float().item(),
Expand Down
Loading
Loading