Repository navigation
laya: add native CUDA inference and optimize batch-1 latency - #90
Conversation
hsliuustc0106
left a comment
There was a problem hiding this comment.
Reviewed commit c2792b1. No actionable correctness findings in the inspected preprocessing, executor ownership, Graph reuse, decoding, HTTP, and Qwen JSON paths.
Independent validation passed: 83 workspace Rust tests, two additional cached-checkpoint tests, formatting, Clippy, release builds, both CUDA bundle builds, eight build-wrapper tests, eight benchmark tests, smoke-fixture validation, and strict documentation build. The remaining eight opt-in tests were not run. Twenty requests matched upstream Laya 0.3.20 token IDs, marker positions, question types, and usage. Eighteen requests per mode matched response JSON and bitwise FP32 raw heads across original Graph, optimized Graph, and optimized eager execution, including changing marker counts within a reused shape, 16 mixed long questions, and empty questions. Real worker/frontend HTTP acceptance passed with 20 exact responses, four expected rejections, and four health checks; the first inference after readiness was verified.
The matched A/B comparison used the same frozen native CLI and pinned checkpoint, batch/concurrency 1, Graph replay and workspace reuse on both sides, and original versus optimized RoPE/GEMM kernels. After separate feasibility checks, each configuration had two measured passes with 20 warmups and 100 requests per input, in A1/B1/B2/A2 order. Each table row pools 200 samples; both passes showed improvement.
| Input | Original p50 | Optimized p50 | Reduction |
|---|---|---|---|
| Choice | 2.595 ms | 1.660 ms | 36.0% |
| Score | 2.643 ms | 1.686 ms | 36.2% |
| Short | 2.608 ms | 1.672 ms | 35.9% |
| Medium | 2.997 ms | 2.020 ms | 32.6% |
| Long | 4.132 ms | 3.124 ms | 24.4% |
GPU 2 was exclusively reserved for these runs. CUDA reports NVIDIA L20X, compute capability 9.0, 143166 MiB; the scheduler labels the host GPUs H200. Other GPUs on the host were active and clocks were not locked. Timing includes native request processing, GPU execution, readback, and serialization; it excludes startup, cold Graph capture, pipe waits, and HTTP. These repetitions describe observed variability, not confidence intervals. Upstream model-forward parity and inference-failure injection were not rerun; this does not establish upstream Python or HTTP performance. Task-owned GPU resources were released.
Purpose
Add native English Laya inference to the Rust frontend: preprocessing, encoder/decision layers, scorer/action heads and answer decoding. On one H800, the matched batch-1 native CLI comparison reduces warm p50 latency by 24.9%–37.6%.
SerialScheduler. Inference runs in Rust/CUDA; Python handles export and reference measurement.Graphs are cached by padded
(B, L). Marker/head sizes can vary within that shape, so gathering and heads remain eager, followed by synchronous readback for CPU decoding.Test Plan
System1-Omni Version / Commit: base
f4ee2127925766969becd6d87fd6224a74d56c2a; headc2792b1ba5810ca1bdd96200b59c50c3fbc8c0da. Runtime code is unchanged from validated commit7b87adcbec0fd4e47779f000f65a9ca696fcbda1; the latest-main merge only updates documentation.Use the same frozen CLI, checkpoint, inputs and H800, batch/concurrency 1. Configurations run in forward/reverse order, with two passes of 100 measured requests per input after 20 warmups. Functional checks, loading and warmup are excluded. The GPU was shared; clocks were not locked.
Test Result
Complete native CLI, original RoPE/GEMM → optimized configuration. Both sides use Graph replay and reusable workspaces. Timing includes parsing, tokenization, uploads, inference, readback and response serialization; HTTP and cold Graph capture are excluded. These measurements precede current-main integration; the integrated worker has not been retimed.
Latencies are in ms; each pair is original → optimized. The ratio is optimized/original p50.
Percentiles use linear interpolation at
q × (N−1), matching the original benchmark; this table pools 200 samples. Engine req/s is1000 / mean(engine_wall_ms), excluding CLI pipe waits and HTTP.Two measured passes and observed variability
Each cell is run 1 / run 2; latencies are in ms. These repetitions describe observed variability, not confidence intervals.
Separate optimized eager→Graph runs reduce p99 by 10.6%–17.9%. Workspace gains were not isolated; these wall-clock samples alone cannot attribute tail changes to launch or allocation overhead.
Upstream reference requested in #49: the published H800 comparison covers short/long, multiple questions and 16 distinct rows padded to L=512, with one feasibility pass, two interleaved measured passes, per-shape p50/p95, req/s, Rust/upstream ratios and separate CUDA-event timings. Eager encoder+decision latency remained close to Laya 0.3.20: ratios 0.984–1.005; all 30 final-hidden comparisons had zero error. Its host-boundary differences are documented. A matched upstream full-request or HTTP speedup has not been measured.
Validation: 83 Rust tests, fmt, Clippy, release builds, CUDA/benchmark tests and strict docs build passed; 10 asset/GPU tests were skipped locally. Fresh Linux CLI/worker/frontend builds and independent code review passed. On H800, 24 CLI and 24 HTTP responses matched the pinned Python JSON reference; raw CLI heads matched the frozen native reference bitwise. Four rejection and four health checks passed, and resources were released. Existing raw-head differences from Python are documented.
Demo / evidence
Model:
convaiinnovations/laya@55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851. FP16 embeddings, BF16 linear layers with FP32 accumulation, FP32 residuals/normalization.Build/run, environment and reproduction · Validation and frozen binary/library hashes · All raw samples and integration results
Self-review