Skip to content

laya: add native CUDA inference and optimize batch-1 latency - #90

Merged
hsliuustc0106 merged 5 commits into
ThinkFlowLab:mainfrom
linear3735:codex/laya-batch1-cuda
Oct 5, 2026
Merged

hsliuustc0106 merged 5 commits into
ThinkFlowLab:mainfrom
linear3735:codex/laya-batch1-cuda

Conversation

@linear3735

Copy link
Copy Markdown

Purpose

Add native English Laya inference to the Rust frontend: preprocessing, encoder/decision layers, scorer/action heads and answer decoding. On one H800, the matched batch-1 native CLI comparison reduces warm p50 latency by 24.9%–37.6%.

  • Optimize RoPE and shape-specific GEMMs; reuse GPU workspaces and encoder/decision CUDA Graphs.
  • Use the existing frontend and SerialScheduler. Inference runs in Rust/CUDA; Python handles export and reference measurement.
  • Preserve Qwen JSON number and duplicate-key handling when Laya enables arbitrary-precision JSON.

Graphs are cached by padded (B, L). Marker/head sizes can vary within that shape, so gathering and heads remain eager, followed by synchronous readback for CPU decoding.

Test Plan

System1-Omni Version / Commit: base f4ee2127925766969becd6d87fd6224a74d56c2a; head c2792b1ba5810ca1bdd96200b59c50c3fbc8c0da. Runtime code is unchanged from validated commit 7b87adcbec0fd4e47779f000f65a9ca696fcbda1; the latest-main merge only updates documentation.

cargo fmt --all --check
cargo clippy --workspace --locked --all-targets --features omni-laya/serve -- -D warnings
cargo test --workspace --locked --features omni-laya/serve
cargo build --workspace --release --locked --features omni-laya/serve
python3 -m unittest discover -s tests/cuda -p 'test_*.py'
python3 benchmarks/bench.py validate benchmarks/smoke.jsonl
python3 -m unittest discover -s tests/benchmarks -p 'test_*.py'
mkdocs build --strict

Use the same frozen CLI, checkpoint, inputs and H800, batch/concurrency 1. Configurations run in forward/reverse order, with two passes of 100 measured requests per input after 20 warmups. Functional checks, loading and warmup are excluded. The GPU was shared; clocks were not locked.

Test Result

Complete native CLI, original RoPE/GEMM → optimized configuration. Both sides use Graph replay and reusable workspaces. Timing includes parsing, tokenization, uploads, inference, readback and response serialization; HTTP and cold Graph capture are excluded. These measurements precede current-main integration; the integrated worker has not been retimed.

Latencies are in ms; each pair is original → optimized. The ratio is optimized/original p50.

Input p50 p90 p95 p99 Engine req/s Ratio
Choice, L48 2.543 → 1.596 2.555 → 1.617 2.558 → 1.633 2.562 → 1.666 393.2 → 624.5 0.628
Score, L64 2.582 → 1.610 2.590 → 1.620 2.596 → 1.625 2.604 → 1.629 387.2 → 620.7 0.624
Short, L48 2.542 → 1.596 2.548 → 1.602 2.550 → 1.603 2.556 → 1.606 393.4 → 626.7 0.628
Medium, L176 2.867 → 1.921 2.886 → 1.935 2.888 → 1.942 2.896 → 1.964 348.9 → 520.5 0.670
Long, L512 3.993 → 3.000 4.081 → 3.026 4.090 → 3.033 4.315 → 3.050 249.1 → 333.3 0.751

Percentiles use linear interpolation at q × (N−1), matching the original benchmark; this table pools 200 samples. Engine req/s is 1000 / mean(engine_wall_ms), excluding CLI pipe waits and HTTP.

Two measured passes and observed variability

Each cell is run 1 / run 2; latencies are in ms. These repetitions describe observed variability, not confidence intervals.

Input Original p50 Original p95 Optimized p50 Optimized p95 Original req/s Optimized req/s Ratio
Choice, L48 2.538 / 2.544 2.559 / 2.557 1.597 / 1.593 1.613 / 1.637 393.5 / 392.9 624.5 / 624.4 0.629 / 0.626
Score, L64 2.585 / 2.579 2.600 / 2.588 1.612 / 1.607 1.625 / 1.621 386.7 / 387.7 619.9 / 621.5 0.624 / 0.623
Short, L48 2.542 / 2.541 2.549 / 2.550 1.598 / 1.594 1.605 / 1.600 393.3 / 393.4 625.9 / 627.5 0.629 / 0.627
Medium, L176 2.857 / 2.875 2.883 / 2.891 1.919 / 1.922 1.942 / 1.940 350.0 / 347.7 520.8 / 520.3 0.672 / 0.668
Long, L512 4.005 / 3.981 4.095 / 4.036 3.014 / 2.991 3.046 / 3.013 248.1 / 250.0 331.9 / 334.8 0.752 / 0.751

Separate optimized eager→Graph runs reduce p99 by 10.6%–17.9%. Workspace gains were not isolated; these wall-clock samples alone cannot attribute tail changes to launch or allocation overhead.

Upstream reference requested in #49: the published H800 comparison covers short/long, multiple questions and 16 distinct rows padded to L=512, with one feasibility pass, two interleaved measured passes, per-shape p50/p95, req/s, Rust/upstream ratios and separate CUDA-event timings. Eager encoder+decision latency remained close to Laya 0.3.20: ratios 0.984–1.005; all 30 final-hidden comparisons had zero error. Its host-boundary differences are documented. A matched upstream full-request or HTTP speedup has not been measured.

Validation: 83 Rust tests, fmt, Clippy, release builds, CUDA/benchmark tests and strict docs build passed; 10 asset/GPU tests were skipped locally. Fresh Linux CLI/worker/frontend builds and independent code review passed. On H800, 24 CLI and 24 HTTP responses matched the pinned Python JSON reference; raw CLI heads matched the frozen native reference bitwise. Four rejection and four health checks passed, and resources were released. Existing raw-head differences from Python are documented.

Demo / evidence

Model: convaiinnovations/laya@55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851. FP16 embeddings, BF16 linear layers with FP32 accumulation, FP32 residuals/normalization.

Build/run, environment and reproduction · Validation and frozen binary/library hashes · All raw samples and integration results

Self-review

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed commit c2792b1. No actionable correctness findings in the inspected preprocessing, executor ownership, Graph reuse, decoding, HTTP, and Qwen JSON paths.

Independent validation passed: 83 workspace Rust tests, two additional cached-checkpoint tests, formatting, Clippy, release builds, both CUDA bundle builds, eight build-wrapper tests, eight benchmark tests, smoke-fixture validation, and strict documentation build. The remaining eight opt-in tests were not run. Twenty requests matched upstream Laya 0.3.20 token IDs, marker positions, question types, and usage. Eighteen requests per mode matched response JSON and bitwise FP32 raw heads across original Graph, optimized Graph, and optimized eager execution, including changing marker counts within a reused shape, 16 mixed long questions, and empty questions. Real worker/frontend HTTP acceptance passed with 20 exact responses, four expected rejections, and four health checks; the first inference after readiness was verified.

The matched A/B comparison used the same frozen native CLI and pinned checkpoint, batch/concurrency 1, Graph replay and workspace reuse on both sides, and original versus optimized RoPE/GEMM kernels. After separate feasibility checks, each configuration had two measured passes with 20 warmups and 100 requests per input, in A1/B1/B2/A2 order. Each table row pools 200 samples; both passes showed improvement.

Input Original p50 Optimized p50 Reduction
Choice 2.595 ms 1.660 ms 36.0%
Score 2.643 ms 1.686 ms 36.2%
Short 2.608 ms 1.672 ms 35.9%
Medium 2.997 ms 2.020 ms 32.6%
Long 4.132 ms 3.124 ms 24.4%

GPU 2 was exclusively reserved for these runs. CUDA reports NVIDIA L20X, compute capability 9.0, 143166 MiB; the scheduler labels the host GPUs H200. Other GPUs on the host were active and clocks were not locked. Timing includes native request processing, GPU execution, readback, and serialization; it excludes startup, cold Graph capture, pipe waits, and HTTP. These repetitions describe observed variability, not confidence intervals. Upstream model-forward parity and inference-failure injection were not rerun; this does not establish upstream Python or HTTP performance. Task-owned GPU resources were released.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants