Skip to content

Latest commit

 

History

281 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

System1-Omni

Documentation site Rust stable toolchain CUDA compute capability 8.0 or newer Contributions welcome

Fast decision-model serving for agents.

How It Works · First Decision · Demo · Supported Models · 中文 · Documentation · Benchmarks · Roadmap · Contributing

About

System1-Omni serves decision models for agents: route a support ticket to billing, score its urgency, or decide whether it asks for a refund. Models return structured decisions from a prefill pass. Start with the complete CPU walkthrough, inspect the real request and response, and check the model and hardware matrix.

The community-maintained engine combines a Rust frontend with model-owned execution and CUDA backends. The target architecture separates processing and scheduling from model execution; a native Metal backend is planned, while LAYA already has a Python worker for Apple GPUs through PyTorch MPS.

The Rust frontend forwards requests to a separately running model worker. The Cua-S1 4B 0.2 text adapter and Open-Jev-27B-v1.1 have native workers using shared CUDA kernels in this repository.

News

Features

  • Rust serving frontend. API, request lifecycle, and response delivery through a small engine interface, forwarding requests to separately running model workers.
  • Model-owned execution. Model executors own weights, forward passes, learned heads, device state, and kernel selection. Native workers have separate processing and executor modules, with shared FIFO admission and blocking dispatch in the native runtime.
  • Native CUDA workers. The Cua-S1 4B 0.2 text adapter and Open-Jev-27B-v1.1 run as native workers with shared CUDA kernels. Cua-S1 also has a Python worker that serves as the correctness reference.
  • LAYA text serving. LAYA runs as an external CPU Python worker, the in-repository Python MPS/CPU worker, or a native Rust/CUDA worker on Hopper.
  • CUDA backend and planned Metal backend. High-performance GPU operations for NVIDIA GPUs, with a native Apple-GPU backend planned alongside it.
  • Benchmark harness. Request replay, output-fidelity checks, and a CUDA comparison protocol across serving backends.

How It Works

Share processing and scheduling; let each model own its execution.

System1-Omni target architecture: Rust frontend, independent processing and batching layers, model executors, and CUDA and Metal backends

The diagram shows the target architecture. Today the frontend forwards HTTP requests to separately running workers, whose handlers coordinate independent processors and executors. Native workers use shared FIFO admission and blocking dispatch per loaded executor. Processing orchestration, batch budgets, compatibility grouping and dynamic batching remain planned.

Layer Responsibility Native target implementation
Rust frontend API transport, request forwarding, and response delivery. Rust.
Processing layer Independent pre/postprocessing modules with model-specific processors for input preparation and output interpretation. Rust CPU processing; GPU transforms use backends.
Scheduler / batcher Queue admission, batch budgets, compatibility grouping, batch assembly, request bookkeeping, and result routing. Rust host policy; GPU packing uses backends.
Model executors Weights, forward passes, learned heads, device state, and kernel selection. Rust orchestration calling backend operations.
CUDA backend High-performance GPU operations for NVIDIA GPUs. Rust bindings/dispatch and CUDA C++ kernels.
Metal backend (planned) High-performance GPU operations for Apple GPUs. Rust bindings/dispatch and Metal shaders.

In the target design, the shared worker runtime invokes processors, schedules compatible work, calls the model executor, and routes each output back to its request. Tokenization, modality transforms, and response interpretation remain model-specific plugins, separate from the forward implementation. Models declare batch constraints; batch adapters pack inputs and unpack outputs using the executor's supported layout. A shared scheduler must not batch incompatible models or inputs, and dynamic batching requires executor support for real batches.

These are logical layers: the runtime and executor can share a worker process. Request bookkeeping belongs to the runtime; model device state belongs to the executor. Backends can optimize for their hardware without requiring identical internal implementations.

The architecture and integration contracts define processor/executor boundaries, compatibility grouping, state and buffer lifetimes, and result reconstruction. They also distinguish the native target from current single-prompt execution and CPU decision heads.

Repository Layout

Implementation code lives under src/; recipes and documentation stay at the repository root.

Directory Responsibility
src/frontend/ Rust serving code, Python worker adapters, and the small engine interface.
src/runtime/ Shared native execution admission and blocking dispatch.
src/models/ Model contracts, existing worker pipelines, and model executors, including the shared Qwen3.5/3.8 prefill implementation.
src/backends/cuda/ NVIDIA GPU operations and kernel integration.
src/backends/metal/ Apple GPU operations and kernel integration.
recipe/ Model setup instructions, launch commands, configuration examples, and example requests.
docs/ Project documentation and architecture assets.

The frontend, native runtime, the native workers, their shared Qwen3.5/3.8 prefill implementation and the Laya CUDA backend are Cargo workspace members. The other model and backend directories currently document planned work; they do not prescribe process boundaries.

Getting Started

Follow Your first decision on CPU (中文入口) for prerequisites, pinned dependencies, worker and frontend startup, readiness checks, a refund request and expected output. The recommended path uses the upstream LAYA English worker on Linux x86_64 with Python 3.12 and stable Rust; it needs no GPU or weight export.

The recorded CPU demo includes the actual JSON and checks for choice, score and noul. Download/build/loading time is separate from inference; no setup-duration or CPU speed claim is made. For other hardware, see the LAYA MPS recipe or the accelerated native Open-Jev recipe. The frontend documentation describes transport and configuration.

Supported Models

LAYA text serving uses the upstream CPU worker, the in-repository Python MPS/CPU worker, or a native Rust/CUDA worker on Hopper. The Cua-S1 4B 0.2 text adapter runs as a Python worker or as a native worker on CUDA. Open-Jev-27B-v1.1 runs as a native Rust/CUDA worker. Cua-S1 also has a Python screenshot worker, and CLM has a stub-encoder contract recipe:

Model Status
LAYA External worker; Python worker on Apple Silicon (MPS) and CPU; CPU checkpoint reader; native Rust/CUDA worker on Hopper
Cua-S1 4B 0.2 (text adapter) Python worker; native worker, CUDA, run on sm_89
Cua-S1 4B 0.2 (multimodal adapter) Python CUDA worker; one PNG/JPEG screenshot, choice; native screenshot execution remains in progress
Open-Jev-27B-v1.1 Native Rust/CUDA worker; eager independent text candidates; H200 validation
CLM-v0.1-8B External worker with a CPU stub encoder; contract checks only, real Qwen3-8B decisions unverified by this recipe

Supported models and hardware lists the devices and where each worker has been run.

Benchmarks

See the GPU serving benchmark for request replay, output-fidelity checks, and the CUDA comparison protocol. The Open-Jev H200 results cover 74 single-candidate requests and a matched comparison with raw HF Transformers and OpenJev-Fast. The Open-Jev optimization notes record PR-by-PR Rust/CUDA changes and isolated A/B measurements, with figures, numerical checks and links to the separate experiments.

Roadmap

The current focus is the native Cua-S1 and Open-Jev CUDA workers and the serving benchmark harness. Planned work extends the shared runtime with processing orchestration, admission budgets, compatibility grouping and bounded dynamic batching with batch-capable executors, additional model engines and GPU backends including Metal, and per-model performance measurements as implementations are added and validated.

🤝 Contributing

System1-Omni is open to contributions across serving, models, backends, benchmarks, and documentation. A reproducible bug report, a carefully measured benchmark, or a clearer recipe can be just as useful as a kernel optimization.

Review the contributing guide before opening a pull request: self-review the full diff, keep serving, processing, scheduling, and model execution separate according to the architecture contracts, and run the checks appropriate to your changes.

Have an idea or found a problem? Open an issue with the details, or send a pull request. For larger changes, start a discussion in an issue so we can work through the design together.

Stay Tuned with Us

If you find system1-omni useful, give us a star on GitHub to support the project and help others discover it!

GitHub repository screenshot demonstrating a click on Star, turning the star yellow and showing Starred

About

community maintained vllm/vllm-omni style inference engine for system1 models in Rust for extreme efficient performance

Resources

Contributing

Stars

15 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages