Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
82 changes: 80 additions & 2 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
[workspace]
members = ["src/frontend", "src/runtime", "src/models/cua_s1/native", "src/models/qwen3_5/native", "src/models/open_jev/native", "src/models/laya"]
members = ["src/frontend", "src/runtime", "src/models/cua_s1/native", "src/models/qwen3_5/native", "src/models/open_jev/native", "src/models/laya", "src/backends/cuda"]
resolver = "3"
18 changes: 9 additions & 9 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,8 +67,8 @@ shared CUDA kernels in this repository.
- **Native CUDA workers.** The Cua-S1 4B 0.2 `text` adapter and
Open-Jev-27B-v1.1 run as native workers with shared CUDA kernels. Cua-S1
also has a Python worker that serves as the correctness reference.
- **LAYA text serving.** LAYA runs as an external CPU Python worker or the
in-repository Python MPS/CPU worker, with a separate native CPU checkpoint reader.
- **LAYA text serving.** LAYA runs as an external CPU Python worker, the
in-repository Python MPS/CPU worker, or a native Rust/CUDA worker on Hopper.
- **CUDA backend and planned Metal backend.** High-performance GPU operations
for NVIDIA GPUs, with a native Apple-GPU backend planned alongside it.
- **Benchmark harness.** Request replay, output-fidelity checks, and a CUDA
Expand All @@ -82,7 +82,7 @@ Share processing and scheduling; let each model own its execution.

The diagram shows the **target architecture**. Today the frontend forwards HTTP
requests to separately running workers, whose handlers coordinate independent
processors and executors. Both native workers use shared FIFO admission and
processors and executors. Native workers use shared FIFO admission and
blocking dispatch per loaded executor. Processing orchestration, batch budgets,
compatibility grouping and dynamic batching remain planned.

Expand Down Expand Up @@ -128,8 +128,8 @@ repository root.
| [`recipe/`](recipe/) | Model setup instructions, launch commands, configuration examples, and example requests. |
| [`docs/`](docs/) | Project documentation and architecture assets. |

The frontend, native runtime, both native workers, their shared Qwen3.5/3.8
prefill implementation and the Laya checkpoint reader are Cargo workspace members.
The frontend, native runtime, the native workers, their shared Qwen3.5/3.8
prefill implementation and the Laya CUDA backend are Cargo workspace members.
The other model and backend directories currently document planned work;
they do not prescribe process boundaries.

Expand All @@ -150,15 +150,15 @@ The [frontend documentation](src/frontend/README.md) describes transport and con

## Supported Models

LAYA text serving uses the upstream CPU worker or the in-repository Python
MPS/CPU worker; native Rust model execution is still planned. The Cua-S1 4B 0.2 `text` adapter
LAYA text serving uses the upstream CPU worker, the in-repository Python MPS/CPU
worker, or a native Rust/CUDA worker on Hopper. The Cua-S1 4B 0.2 `text` adapter
runs as a Python worker or as a native worker on CUDA. Open-Jev-27B-v1.1
runs as a native Rust/CUDA worker. Cua-S1 also has a Python screenshot worker,
and CLM has a stub-encoder contract recipe:

| Model | Status |
| --- | --- |
| LAYA | [External worker](recipe/laya/README.md); [Python worker on Apple Silicon (MPS) and CPU](recipe/laya/apple-silicon.md); [CPU checkpoint reader](src/models/laya/README.md); native Rust execution planned |
| LAYA | [External worker](recipe/laya/README.md); [Python worker on Apple Silicon (MPS) and CPU](recipe/laya/apple-silicon.md); [CPU checkpoint reader](src/models/laya/README.md); [native Rust/CUDA worker on Hopper](recipe/laya/native/README.md) |
| Cua-S1 4B 0.2 (`text` adapter) | [Python worker](recipe/cua_s1/text.md); [native worker](recipe/cua_s1/native.md), CUDA, run on sm_89 |
| Cua-S1 4B 0.2 (`multimodal` adapter) | [Python CUDA worker](src/frontend/cua_s1.py); one PNG/JPEG screenshot, `choice`; native screenshot execution remains in progress |
| Open-Jev-27B-v1.1 | [Native Rust/CUDA worker](recipe/open_jev/native.md); eager independent text candidates; [H200 validation](recipe/open_jev/validation.md) |
Expand All @@ -183,7 +183,7 @@ The current focus is the native Cua-S1 and Open-Jev CUDA workers and the serving
benchmark harness. Planned work extends the shared runtime with processing
orchestration, admission budgets, compatibility grouping and bounded dynamic
batching with batch-capable executors,
the in-repository LAYA model engine, additional model engines and GPU backends
additional model engines and GPU backends
including Metal, and per-model
performance measurements as implementations are added and validated.

Expand Down
44 changes: 28 additions & 16 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,23 +8,25 @@ design. Concrete input/output types follow each executor's supported layout.

The [Rust frontend](../src/frontend/README.md) currently forwards HTTP requests
to separately running workers. Cua-S1 and Open-Jev have native Rust/CUDA workers
that share the [Qwen3.5/3.8 executor](../src/models/qwen3_5/native/). Their
model-specific workers coordinate independent processor and executor modules
through `prepare` → `execute` → `finish`. The shared Qwen executor accepts one
prompt per forward call. Both workers use the
that share the [Qwen3.5/3.8 executor](../src/models/qwen3_5/native/), which accepts
one prompt per forward call. [Laya's native worker](../src/models/laya/README.md)
uses a separate Hopper CUDA backend for one complete padded request. All three
coordinate independent processors and executors through
`prepare` → `execute` → `finish` and use the
[native runtime](../src/runtime/README.md) for FIFO admission and blocking dispatch
per loaded executor. Shared processing orchestration, batch budgets,
compatibility grouping and dynamic batching are planned.

The native workers currently compute their decision heads on the CPU after
downloading the final hidden state. GPU head execution belongs to the target
model/backend integration. LAYA's native executor and the Metal backend are
also planned; Python workers retain their documented reference/serving roles.
Qwen workers compute their decision heads on the CPU after downloading the final
hidden state. Laya computes its scorer and action head on CUDA. Its fixed-shape
Graph captures Encoder/Decision; gather, scorer/action head and synchronized
readback remain outside capture. Native Metal remains planned; Python workers
retain their documented reference/serving roles.

## Native worker boundaries

Both native workers separate `processing.rs` from `executor.rs`; `engine.rs`
assembles them with a `SerialScheduler` per loaded executor, and the HTTP handler
The native workers separate `processing.rs` from `executor.rs`; their worker
assembly owns a `SerialScheduler` per loaded executor, and the HTTP handler
coordinates the three stages. Preparation validates the entire request before
any forward call and returns executor inputs plus a response context. The context retains question
and candidate identity, usage, and response metadata outside the executor.
Expand All @@ -33,13 +35,21 @@ and candidate identity, usage, and response metadata outside the executor.
| --- | --- | --- | --- |
| Cua-S1 | One unpadded token-ID vector and option count per question, in request order. | One FP32 answer-letter logit vector per question. | Per-question softmax, choice/confidence, ordered answers, and token usage. |
| Open-Jev | Token-ID vectors grouped by question, then independent candidate, in request order. | One FP32 learned scalar per candidate in the same grouping. | Add the `noul` false logit of zero, calibrate across each complete question, and restore typed answers, usage, and metadata. |
| Laya | One padded request: token IDs, true lengths, question types and ordered option markers; at most 16 questions, 512 tokens per row and 2048 markers. | Per-question FP32 option logits and two action logits copied back after GPU heads. | Calibrate and decode ordered `choice`, `score` and `noul` answers, usage and metadata. |

These input collections are serial work, not GPU batches. Shared runtime
Qwen input collections are serial work, not GPU batches; Laya batches questions
within one request. Shared runtime
admission precedes blocking dispatch: Cua-S1 admits one question forward at a
time; Open-Jev admits one complete request. Cua-S1's CPU letter projection stays
outside admission; Open-Jev's scalar heads remain inside its request unit. The
model mutexes guard mutable state, retaining per-question/request granularity.
Executors own the loaded Qwen model and CPU head weights, preserving FP64 accumulation and
time; Open-Jev and Laya admit one complete request. Cua-S1's CPU letter projection
stays outside admission; Open-Jev's scalar heads and Laya's GPU heads and
synchronized readback remain inside their request unit. Qwen model mutexes guard
mutable state. Laya's dedicated owning thread confines its non-Send CUDA state
and receives admitted work over a rendezvous channel. Cancellation after
dispatch retains the scheduler permit until execution completes. Laya
synchronizes and disposes a failed model before returning an inference error
and reports unavailable health thereafter.

Qwen executors own loaded models and CPU head weights, preserving FP64 accumulation and
the existing FP32 rounding and bias order. Finishing checks output cardinality
before reconstruction. HTTP validation, error status/body conventions, and real
warmup before readiness remain model-specific and unchanged.
Expand Down Expand Up @@ -81,7 +91,9 @@ semantics.

Cua-S1 prepares one prompt per question and reads option-letter logits.
Open-Jev prepares independent candidate prompts and normalizes across the
complete question's candidates. A request, question, and GPU batch therefore
complete question's candidates. Laya pads prepared questions into one request
batch and normalizes each question's complete option set. A request, question,
and GPU batch therefore
have different boundaries. Scheduler grouping must preserve those distinctions;
probabilities must not be normalized across unrelated questions or requests.

Expand Down
9 changes: 5 additions & 4 deletions docs/supported-models.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@ Models that are being added are also tracked in issues labeled [new model](https
| --- | --- | --- | --- | --- | --- |
| LAYA, English checkpoint | [External worker](../recipe/laya/README.md) running the upstream Laya runtime, for text requests | Validated ([#2](https://github.com/ThinkFlowLab/system1-omni/pull/2)); [CPU demo](../recipe/laya/validation.md) | Unverified ([#39](https://github.com/ThinkFlowLab/system1-omni/issues/39)) | Use the dedicated MPS worker below | Python 3.12, [pinned CPU dependencies](../recipe/laya/requirements-cpu.txt) |
| LAYA, English checkpoint | [Python MPS/CPU worker](../recipe/laya/apple-silicon.md) in this repository | Contract checks documented ([#67](https://github.com/ThinkFlowLab/system1-omni/pull/67)) | No validation recorded for this path | **PyTorch MPS validated** on M1 Pro; M4/M5 checks documented ([#30](https://github.com/ThinkFlowLab/system1-omni/pull/30), [#67](https://github.com/ThinkFlowLab/system1-omni/pull/67)) | Python 3.12, [MPS dependency versions](../recipe/laya/requirements-mps.txt); native Metal execution remains planned |
| LAYA | Native Rust model execution | Planned ([#14](https://github.com/ThinkFlowLab/system1-omni/issues/14)) | Planned ([#14](https://github.com/ThinkFlowLab/system1-omni/issues/14)) | Planned ([#3](https://github.com/ThinkFlowLab/system1-omni/issues/3)) | |
| LAYA, English checkpoint | [Native Rust worker](../recipe/laya/native/README.md) | Processing/checkpoint checks; no CPU inference | Hopper `sm_90a`; [validation scope](../recipe/laya/native/VALIDATION.md) | Not supported | CUDA toolkit, TileLang for AOT generation, pinned checkpoint and rotary tables |
| Cua-S1 4B 0.2, `text` adapter | [Reference worker](../recipe/cua_s1/text.md) on Transformers and PEFT | Unverified | Validated ([#13](https://github.com/ThinkFlowLab/system1-omni/pull/13)) | Unverified | Python 3.12, the versions in `requirements-text.txt` |
| Cua-S1 4B 0.2, `text` adapter | [Native Rust worker](../recipe/cua_s1/native.md) on the [Qwen3.5 CUDA kernels](../src/backends/cuda/qwen3_5/README.md) | Not supported | Validated on compute capability 8.9 ([#19](https://github.com/ThinkFlowLab/system1-omni/pull/19), [#52](https://github.com/ThinkFlowLab/system1-omni/pull/52)) | Not supported | Compute capability 8.0 or newer, the CUDA toolkit to build, weights merged with `export_text_merged.py` |
| Cua-S1 4B 0.2, `multimodal` adapter | Reference worker on Transformers and PEFT, [`src/frontend/cua_s1.py`](../src/frontend/cua_s1.py); no recipe yet | Not supported | Validated ([#17](https://github.com/ThinkFlowLab/system1-omni/pull/17), [#18](https://github.com/ThinkFlowLab/system1-omni/pull/18)) | Not supported | The state is one PNG or JPEG image; upstream's `weights.lock.json` next to the base weights |
Expand All @@ -28,7 +28,8 @@ not validate decision quality. MPS validation above is for a Python/PyTorch
worker, not a native Metal backend.

The [architecture contracts](architecture.md) describe the native target.
Shared processing orchestration, scheduling, dynamic batching, and GPU decision
heads are planned; the existing native workers run independent single-prompt
prefills and compute their heads on the CPU. These target layers do not expand
Shared processing orchestration and dynamic batching remain planned. Native
workers reuse serial admission. Qwen workers run independent single-prompt
prefills with CPU heads; Laya packs questions within one request and runs its
scorer/action head on CUDA. These target layers do not expand
the validated model or hardware coverage above.
2 changes: 2 additions & 0 deletions mkdocs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -59,6 +59,8 @@ nav:
- Recipes:
- recipe/README.md
- Laya text worker: recipe/laya/README.md
- Laya native CUDA worker: recipe/laya/native/README.md
- Laya CUDA validation: recipe/laya/native/VALIDATION.md
- Laya CPU reproduction and demo: recipe/laya/validation.md
- Laya on Apple Silicon: recipe/laya/apple-silicon.md
- Cua-S1 text worker: recipe/cua_s1/text.md
Expand Down
2 changes: 2 additions & 0 deletions recipe/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,8 @@ For a first real decision, follow the [complete CPU walkthrough](../docs/getting

- [Laya text worker](laya/README.md): start the external Python worker, connect the
Rust frontend and compare direct and proxied responses.
- [Laya native CUDA worker](laya/native/README.md): build the Hopper bundle and
serve English text decisions with Rust and CUDA.
- [Laya on Apple Silicon](laya/apple-silicon.md): serve Laya on the Mac GPU with the Laya
worker, put the frontend in front of it and run the benchmarks.
- [Cua-S1 4B 0.2 text worker](cua_s1/text.md): download the pinned weights, start
Expand Down
1 change: 1 addition & 0 deletions recipe/laya/README.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
# Laya text worker

This recipe runs the external Laya Python package behind the Rust frontend.
For the Rust/CUDA worker on Hopper, see [native CUDA setup](native/README.md).
It validates text decisions; image, audio and video inference are not covered.

Run all commands from the repository root. To serve on the GPU of an Apple Silicon Mac, see
Expand Down
Loading
Loading