diff --git a/README.md b/README.md index dbe457f0..cbcb6333 100644 --- a/README.md +++ b/README.md @@ -86,6 +86,11 @@ processors and executors. Native workers use shared FIFO admission and blocking dispatch per loaded executor. Processing orchestration, batch budgets, compatibility grouping and dynamic batching remain planned. +Open-Jev already packs candidates within one request for selected prefill GEMMs, +with bounded groups and independent sequence state. See the +[native recipe](recipe/open_jev/native.md) and +[matched H200 measurements](benchmarks/prefill_batching/README.md). + | Layer | Responsibility | Native target implementation | | --- | --- | --- | | Rust frontend | API transport, request forwarding, and response delivery. | Rust. | diff --git a/benchmarks/prefill_batching/README.md b/benchmarks/prefill_batching/README.md new file mode 100644 index 00000000..cfe1dbc5 --- /dev/null +++ b/benchmarks/prefill_batching/README.md @@ -0,0 +1,103 @@ +# Open-Jev selective prefill packing on H200 + +Packing input and gate/up projections across candidates reduces mean warm HTTP +latency by **18.04% for four-question requests**, **20.11% for eight-question +requests**, and **10.67% for mixed Choice/Noul/Score examples**. Compared outputs +match the native reference exactly. Single-question mean/p95 increases +**0.21%/1.09%**, within the declared 2% regression limits. + +These are concurrency-1, within-request observations on one H200 with BF16, +from 2026-10-06. Repeated-question workloads are synthetic transformations of +real JevBench cases. Two measured passes show observed variability, not +confidence intervals or production-traffic performance. + +## Implementation + +The adapter preserves candidate order in groups of at most 16 sequences and +4096 tokens. Longer prompts execute alone at their original length. Input and +gate/up GEMMs share packed rows; output/down projections keep per-prompt shapes +to preserve cuBLASLt reduction order. Attention, positions, convolution, GDN +state and final-token readout remain independent. Scores are regrouped by +question before calibration. Runtime admission still covers one whole request. + +CUDA Graph keys contain ordered sequence lengths. Cache misses retain the eager +result; cache hits read freshly embedded inputs. The upstream CUDA ABI5 and +native vision support are preserved. Graph performance is unmeasured here. + +## Matched warm HTTP results + +Each latency cell gives measured pass 1 / 2 means in milliseconds. Reduction +uses the arithmetic mean of those pass means. + +| Workload | Requests/pass | Current main | Packed | Mean reduction | +| --- | ---: | ---: | ---: | ---: | +| Real single-question JevBench | 74 | 48.048 / 48.143 | 48.307 / 48.083 | -0.21% | +| Four repeated questions | 60 | 98.703 / 98.618 | 80.800 / 80.915 | 18.04% | +| Eight repeated questions | 60 | 197.470 / 197.502 | 157.558 / 157.978 | 20.11% | +| Mixed Choice/Noul/Score | 12 | 258.378 / 257.593 | 230.449 / 230.490 | 10.67% | + +The single-question slice contains 74 Noul cases spanning 80–3399 tokens from +[JevBench at f8ce713](https://github.com/fstandhartinger/jevbench/tree/f8ce71361165846101d02ebc83ad44e47ae44fc3). +The four/eight-question slices duplicate the 60 cases of at most 400 tokens +under distinct question IDs. The mixed slice uses the repository's +[three-question fixture](../../tests/open_jev/data/contract.json), appending +0–33 copies of a fixed billing-context sentence across 12 cases. + +The measured passes contain **824 successful requests / 3320 decisions**. +Answers, token usage, model identity and metadata except inference time match +the baseline feasibility reference exactly: maximum probability/Score drift +**0.0**, zero decision flips and zero failures. All declared gates pass: +at least 15% reduction for both repeated-question slices, at most 2% single +mean/p95 regression, drift at most 0.001, zero flips and zero failures. + +## Controls and evidence + +[summary.json](artifacts/20261006/summary.json) records per-pass metrics, gates, +frozen source/checkpoint revisions, source hashes and execution controls. +[timings.csv](artifacts/20261006/timings.csv) retains all 824 measured request +latencies in seconds: one row per case, with baseline/candidate pass 1/2 columns +and the question-answer count for each request. +P95 uses nearest rank within each pass; reported reductions average pass metrics. + +Baseline: `47eff9cdeda01e4847a4fb9634a43f2cab6a233f`. +Measured candidate runtime: `82b5e7e60363405614bccd79ba83ab774e7513f5`; +its source hashes match the final runtime files. Both arms use the same merged +export, tokenizer, trained head, temperature `2.5343690298472983`, max length +16384, CUDA ABI5 library, frontend and 32 MiB GEMM workspace. GPU 2 UUID +`GPU-cbf66259-f4ab-0ede-1811-82037dde5924`, SM90, NUMA 0 / CPUs 0–15 are fixed. +The driver/scheduler labels this H200 device L20X. + +Each arm reuses one worker/frontend pair. Real readiness, validated first +inference, preparation and one feasibility pass per slice are excluded, followed +by exactly two measured passes. Graphs and profiling are disabled during HTTP +measurements; no shared caches are reset or clocks changed. Timing includes +localhost forwarding, tokenization, inference and UTF8 response receipt; +request construction and response JSON parsing are excluded. + +The complete [raw evidence snapshot at 7d03326](https://github.com/ThinkFlowLab/system1-omni/tree/7d03326d51debde070f2a0d3a925405032275270/benchmarks/prefill_batching/artifacts/20261006) +preserves requests, all measured/feasibility responses, collector, frozen plan, +validation logs and JevBench's MIT notice. These files remain in Git history +and a verified local archive; duplicate inputs, process logs and rejected-attempt +records are excluded from the current diff. + +For reproduction, build baseline/candidate worktrees at the pinned runtime SHAs +with matched release options and prepare the pinned export using the +[native recipe](../../recipe/open_jev/native.md). Copy the archived plan and +harness into a fresh run directory and create its `analysis/` directory. Update +host paths, hashes and collector GPU/affinity assertions together, freeze those +controls, then run the collector through the verified scheduler with fixed exact +GPU IDs, NUMA and CPU affinity. Reuse the declared feasibility/two-pass budget. + +## Validation and limits + +Formatting, strict Clippy, locked workspace tests (**90 passed, 13 ignored**), +release build and strict docs pass. Six reserved CUDA reference tests and the +checkpoint packing test pass, covering unequal lengths, reordered shapes, +17-candidate splitting, changed-input cache hits and later singleton execution. +Nsight Systems records **four actual graph launches** in the correctness test; +this trace is outside measured HTTP runs. Task-owned processes exited and the +GPU returned available with 0 MB used. + +Other architectures, higher concurrency, production distributions, peak memory, +cold capture and combined graph performance remain unmeasured. Output fidelity +is against native main; this is not a new full-precision accuracy evaluation. diff --git a/benchmarks/prefill_batching/artifacts/20261006/summary.json b/benchmarks/prefill_batching/artifacts/20261006/summary.json new file mode 100644 index 00000000..c1d50b8d --- /dev/null +++ b/benchmarks/prefill_batching/artifacts/20261006/summary.json @@ -0,0 +1,257 @@ +{ + "date": "2026-10-06", + "baseline_sha": "47eff9cdeda01e4847a4fb9634a43f2cab6a233f", + "candidate_runtime_sha": "82b5e7e60363405614bccd79ba83ab774e7513f5", + "raw_evidence_url": "https://github.com/ThinkFlowLab/system1-omni/tree/7d03326d51debde070f2a0d3a925405032275270/benchmarks/prefill_batching/artifacts/20261006", + "input_manifest_sha256": "02a9a45a381a231fa81579f0f642e5ad23f35410fe5f8e2f013ffac3326a2113", + "collector_sha256": "8ff176a09380aa257b23c9b8a566dd5fa2862ac40280ea76c9c65993cfbd9d4a", + "runtime_source_sha256": { + "src/models/qwen3_5/native/src/model.rs": "161599ea12a890b7f1926fa5ce786ffbfeda3cd53149d87e335b5cede9207aec", + "src/models/open_jev/native/src/executor.rs": "cea8e9f85514d3d4524702f9a3d7c04be159a1bbd028e67c5f187fdea0326888", + "src/models/open_jev/native/src/batching.rs": "32cce2ee7d4ede1f493f6b61fe98efc2e23681aeee018ba614b159364adcb28a", + "src/models/open_jev/native/src/lib.rs": "df6a9d85476ed24e1b5c95acf98735086325d1839b04a7d2e9056531753f54f8" + }, + "controls": { + "independent_variable": "Within-request input and gate/up GEMM packing; output/down projections and mixers stay sequence-local.", + "gpu": "One H200, SM90, 132 SMs; driver/scheduler label L20X", + "gpu_id": 2, + "gpu_uuid": "GPU-cbf66259-f4ab-0ede-1811-82037dde5924", + "numa_node": 0, + "cpu_affinity": "0-15", + "precision": "BF16 weights/activations, FP32 GEMM accumulation", + "cuda_abi": 5, + "cuda_toolkit": "13.0 (nvcc 13.0.88)", + "rustc": "1.98.1", + "model": "Qwen/Qwen3.8-27B with merged Open-Jev-27B-v1.1 adapter and decision head", + "base_revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0", + "checkpoint_revision": "28cf73067d5b337860bbef3c85b8b82ba8730956", + "temperature": 2.5343690298472983, + "max_length": 16384, + "gemm_workspace_bytes": 33554432, + "packing_max_sequences": 16, + "packing_max_tokens": 4096, + "oversized_prompt": "singleton, original length", + "concurrency": 1, + "graphs": "CUA_S1_GRAPH=0 during HTTP measurements; 1 in separate correctness tests", + "profiler": "disabled during measurements", + "matched": "CUDA library, frontend, exported weights, tokenizer, head, temperature and max length", + "cache": "warm; no shared cache resets or clock changes", + "server_reuse": "one worker/frontend pair per arm, reused across all slices", + "budget": { + "feasibility_passes_per_arm_slice": 1, + "measured_passes_per_arm_slice": 2, + "no_extra_measured_runs": true + }, + "timing": "localhost HTTP through the same frontend, includes tokenization, inference and UTF8 response receipt; excludes JSON body construction and JSON parsing, preparation, process-to-readiness, validated first request and full feasibility passes. Profiling disabled during measured runs.", + "p50": "median of per-request latencies within each measured pass", + "p95": "nearest rank within each measured pass", + "reductions": "100 * (1 - mean(candidate pass metrics) / mean(baseline pass metrics))", + "decision_rate": "number of question answers / sum of successful request latencies; excludes gaps between requests", + "acceptance": { + "multi_question_mean_reduction_percent": 15, + "single_mean_max_regression_percent": 2, + "single_p95_max_regression_percent": 2, + "maximum_probability_or_score_drift": 0.001, + "decision_flips": 0, + "request_failures": 0 + }, + "stop_condition": "Stop on reservation/UUID/affinity mismatch, CUDA error, output drift above tolerance, decision flip, request failure or run cap. Preserve failed evidence. Do not silently relax numerical or performance gates." + }, + "workloads": { + "single": { + "requests": 74, + "definition": "original 74 real single-question JevBench Noul requests" + }, + "questions-4": { + "requests": 60, + "definition": "synthetic repeated questions from short JevBench requests" + }, + "questions-8": { + "requests": 60, + "definition": "synthetic repeated questions from short JevBench requests" + }, + "mixed-types": { + "requests": 12, + "definition": "repository Choice/Noul/Score example with varied lengths" + } + }, + "measured_requests": 824, + "measured_decisions": 3320, + "fidelity_comparison": "Every measured and feasibility response matches its baseline feasibility reference after removing metadata.inference_seconds.", + "validation": { + "workspace_tests": { + "passed": 90, + "ignored": 13 + }, + "cuda_reference_tests_passed": 6, + "full_checkpoint_packing_test": "passed with graph replay, updated inputs, ordering, unequal lengths, splitting and singleton execution", + "cuda_graph_launches_in_correctness_trace": 4, + "cuda_graph_trace_sha256": "300dc52f657c7ac976557f868ded9d08ec562f68cf001b69a44e606c870974ea", + "format_clippy_release_build_strict_docs": "passed; original commands and logs in raw evidence snapshot", + "task_owned_process_cleanup": "all exited; GPU returned available with 0 MB used" + }, + "status": "complete", + "results": { + "baseline": { + "single": [ + { + "requests": 74, + "mean_ms": 48.04829193430172, + "p50_ms": 24.96585389599204, + "p95_ms": 221.37836087495089, + "successful_decisions_per_second": 20.812394358728472 + }, + { + "requests": 74, + "mean_ms": 48.14348654267756, + "p50_ms": 24.60308652371168, + "p95_ms": 222.03110624104738, + "successful_decisions_per_second": 20.77124179848367 + } + ], + "questions-4": [ + { + "requests": 60, + "mean_ms": 98.70255975984037, + "p50_ms": 80.30658261850476, + "p95_ms": 154.39793188124895, + "successful_decisions_per_second": 40.52579801104106 + }, + { + "requests": 60, + "mean_ms": 98.6181381624192, + "p50_ms": 80.3435486741364, + "p95_ms": 153.67948170751333, + "successful_decisions_per_second": 40.56048993149919 + } + ], + "questions-8": [ + { + "requests": 60, + "mean_ms": 197.46969395006695, + "p50_ms": 160.7643635943532, + "p95_ms": 310.2688295766711, + "successful_decisions_per_second": 40.51254569738136 + }, + { + "requests": 60, + "mean_ms": 197.50209141833088, + "p50_ms": 160.67188186571002, + "p95_ms": 312.95446306467056, + "successful_decisions_per_second": 40.50590017831827 + } + ], + "mixed-types": [ + { + "requests": 12, + "mean_ms": 258.3781545981765, + "p50_ms": 250.29732752591372, + "p95_ms": 407.54649974405766, + "successful_decisions_per_second": 11.610888717219643 + }, + { + "requests": 12, + "mean_ms": 257.5930858341356, + "p50_ms": 249.6383930556476, + "p95_ms": 405.721134506166, + "successful_decisions_per_second": 11.646275327171251 + } + ] + }, + "candidate": { + "single": [ + { + "requests": 74, + "mean_ms": 48.30653577841617, + "p50_ms": 24.934389162808657, + "p95_ms": 220.3979603946209, + "successful_decisions_per_second": 20.701132546267367 + }, + { + "requests": 74, + "mean_ms": 48.08322312562047, + "p50_ms": 25.28876857832074, + "p95_ms": 227.83073224127293, + "successful_decisions_per_second": 20.797274704057102 + } + ], + "questions-4": [ + { + "requests": 60, + "mean_ms": 80.79964901941518, + "p50_ms": 59.089095797389746, + "p95_ms": 142.11427047848701, + "successful_decisions_per_second": 49.50516553653406 + }, + { + "requests": 60, + "mean_ms": 80.91472735007603, + "p50_ms": 58.53466596454382, + "p95_ms": 142.37547758966684, + "successful_decisions_per_second": 49.43475843024318 + } + ], + "questions-8": [ + { + "requests": 60, + "mean_ms": 157.55837899632752, + "p50_ms": 113.72156580910087, + "p95_ms": 279.7259949147701, + "successful_decisions_per_second": 50.774830580012946 + }, + { + "requests": 60, + "mean_ms": 157.9782015644014, + "p50_ms": 114.35522744432092, + "p95_ms": 280.32094053924084, + "successful_decisions_per_second": 50.63989791489505 + } + ], + "mixed-types": [ + { + "requests": 12, + "mean_ms": 230.44868713865677, + "p50_ms": 231.44325334578753, + "p95_ms": 379.89793717861176, + "successful_decisions_per_second": 13.018082408058826 + }, + { + "requests": 12, + "mean_ms": 230.4904266881446, + "p50_ms": 229.0079789236188, + "p95_ms": 377.71117873489857, + "successful_decisions_per_second": 13.015724961362599 + } + ] + } + }, + "reductions": { + "single": { + "mean_reduction_percent": -0.2058184495515203, + "p95_reduction_percent": -1.0868567040844823 + }, + "questions-4": { + "mean_reduction_percent": 18.044899459455866, + "p95_reduction_percent": 7.656408577908424 + }, + "questions-8": { + "mean_reduction_percent": 20.111614993860428, + "p95_reduction_percent": 10.137034018670422 + }, + "mixed-types": { + "mean_reduction_percent": 10.665735276136978, + "p95_reduction_percent": 6.8438132777811305 + } + }, + "gates": { + "multi_question": true, + "single_mean": true, + "single_p95": true, + "fidelity": true + }, + "maximum_drift": 0.0, + "failures": 0, + "decision_flips": 0, + "dataset_revision": "f8ce71361165846101d02ebc83ad44e47ae44fc3", + "unique_measured_cases": 206 +} diff --git a/benchmarks/prefill_batching/artifacts/20261006/timings.csv b/benchmarks/prefill_batching/artifacts/20261006/timings.csv new file mode 100644 index 00000000..777879ce --- /dev/null +++ b/benchmarks/prefill_batching/artifacts/20261006/timings.csv @@ -0,0 +1,207 @@ +workload,request_id,decisions,baseline_pass_1_s,baseline_pass_2_s,candidate_pass_1_s,candidate_pass_2_s +single,original-policy-01-0,1,0.02081034518778324,0.020888302475214005,0.02070673182606697,0.02102378662675619 +single,original-policy-01-1,1,0.020581857301294804,0.02029157243669033,0.020328269340097904,0.020304175093770027 +single,original-policy-02-0,1,0.020347668789327145,0.020267066545784473,0.020346115343272686,0.020249629393219948 +single,original-policy-02-1,1,0.020271546207368374,0.020191791467368603,0.02025262825191021,0.020153488032519817 +single,original-policy-03-0,1,0.020382529124617577,0.020260430872440338,0.02033732831478119,0.020187209360301495 +single,original-policy-03-1,1,0.020372312515974045,0.020265068858861923,0.02033795416355133,0.02015122678130865 +single,original-policy-04-0,1,0.02021338976919651,0.01998906210064888,0.02011373359709978,0.019818680360913277 +single,original-policy-04-1,1,0.020397797226905823,0.020226269960403442,0.02050396427512169,0.02023342903703451 +single,original-policy-05-0,1,0.020193105563521385,0.02005042787641287,0.020103657618165016,0.019951640628278255 +single,original-policy-05-1,1,0.020449104718863964,0.020225469022989273,0.02030393574386835,0.020195932127535343 +single,original-policy-06-0,1,0.020273881033062935,0.020104880444705486,0.020143773406744003,0.020309917628765106 +single,original-policy-06-1,1,0.02033931389451027,0.020247661508619785,0.020356560125947,0.020288855768740177 +single,original-adequacy-01-0,1,0.02002322394400835,0.02000794466584921,0.02019086852669716,0.019934583455324173 +single,original-adequacy-01-1,1,0.019873691722750664,0.019694486632943153,0.020045402459800243,0.019769652746617794 +single,original-adequacy-02-0,1,0.020181038416922092,0.019852581433951855,0.020120203495025635,0.019890408031642437 +single,original-adequacy-02-1,1,0.02023028116673231,0.02001357264816761,0.020188260823488235,0.019911741837859154 +single,original-adequacy-03-0,1,0.020138688385486603,0.02005852572619915,0.02012857049703598,0.01991879940032959 +single,original-adequacy-03-1,1,0.0201422069221735,0.020051089115440845,0.020174591802060604,0.01998015306890011 +single,original-adequacy-04-0,1,0.020213251933455467,0.02016372699290514,0.020184488967061043,0.01997008826583624 +single,original-adequacy-04-1,1,0.02003743313252926,0.020044841803610325,0.020054891705513,0.01992800273001194 +single,original-adequacy-05-0,1,0.0198267949745059,0.019766349345445633,0.019913638941943645,0.01975338812917471 +single,original-adequacy-05-1,1,0.01978136505931616,0.019727567210793495,0.019874180667102337,0.019759004935622215 +single,original-adequacy-06-0,1,0.020377458073198795,0.020235885865986347,0.020230702124536037,0.02017976902425289 +single,original-adequacy-06-1,1,0.020388481207191944,0.02028157375752926,0.0203027855604887,0.020141683518886566 +single,easy-fact-00,1,0.02024342492222786,0.02012493833899498,0.019986975006759167,0.0198395112529397 +single,easy-fact-01,1,0.020035965368151665,0.019845250993967056,0.019769122824072838,0.019686099141836166 +single,easy-fact-02,1,0.02021977584809065,0.020200642757117748,0.019935132935643196,0.01989852637052536 +single,easy-fact-03,1,0.020341988652944565,0.020295899361371994,0.020071161910891533,0.020047363825142384 +single,easy-fact-04,1,0.019875151105225086,0.019815084524452686,0.019693628884851933,0.019743189215660095 +single,easy-fact-05,1,0.019891698844730854,0.019871583208441734,0.019773351959884167,0.019858533516526222 +single,easy-fact-06,1,0.020053020678460598,0.020104684866964817,0.019894280470907688,0.019848665222525597 +single,easy-fact-07,1,0.019789351150393486,0.019749892875552177,0.019639788195490837,0.019666442647576332 +single,easy-fact-08,1,0.01975057367235422,0.01964633632451296,0.019590686075389385,0.01966224517673254 +single,easy-fact-09,1,0.01967847626656294,0.019676935859024525,0.019604568369686604,0.01983362529426813 +single,easy-fact-10,1,0.01986348908394575,0.019925798289477825,0.019843870773911476,0.020017574541270733 +single,easy-fact-11,1,0.019913519732654095,0.020109256729483604,0.01988764200359583,0.019893267191946507 +single,hard-opus-a-long_policy-11,1,0.21176895778626204,0.21378031093627214,0.21517783030867577,0.21072582714259624 +single,hard-opus-a-long_policy-13,1,0.22137836087495089,0.22203110624104738,0.2203979603946209,0.22783073224127293 +single,hard-opus-a-long_policy-19,1,0.22163983434438705,0.22652530763298273,0.2294031335040927,0.22874982096254826 +single,hard-opus-a-probability-03,1,0.04742553364485502,0.0469998624175787,0.04930546693503857,0.046847143210470676 +single,hard-opus-a-probability-08,1,0.04281137976795435,0.042760965414345264,0.04346830863505602,0.04213203676044941 +single,hard-opus-a-temporal_numeric-03,1,0.049448054283857346,0.0497997272759676,0.050822812132537365,0.049852085299789906 +single,hard-opus-a-temporal_numeric-06,1,0.05815286189317703,0.06036674324423075,0.060041834600269794,0.05882532987743616 +single,hard-opus-a-temporal_numeric-09,1,0.05756783485412598,0.056287022307515144,0.05632245633751154,0.05639549996703863 +single,hard-opus-b-multi_hop-08,1,0.2181630078703165,0.2145283343270421,0.21560550574213266,0.21291996352374554 +single,hard-opus-b-probability-03,1,0.1130249472334981,0.11298779677599669,0.11236789915710688,0.11362842004746199 +single,hard-opus-b-tradeoff-06,1,0.07559008430689573,0.07654078770428896,0.07634077779948711,0.07649533730000257 +single,hard-opus-b-tradeoff-08,1,0.05837902706116438,0.05885340552777052,0.05886319372802973,0.05980461835861206 +single,hard-opus-c-long_policy-04,1,0.3066014423966408,0.3108488405123353,0.31414505932480097,0.31355208065360785 +single,hard-opus-c-long_policy-11,1,0.3298046961426735,0.3276161439716816,0.3304365696385503,0.3229247434064746 +single,hard-opus-c-temporal_numeric-02,1,0.08852543868124485,0.09080446790903807,0.08996532671153545,0.08906852267682552 +single,hard-sol-a-adversarial-06,1,0.02796135190874338,0.02808993775397539,0.029309767298400402,0.027813095599412918 +single,hard-sol-a-adversarial-08,1,0.027125478722155094,0.027529802173376083,0.027244138531386852,0.027020197361707687 +single,hard-sol-a-trap-02,1,0.031932451762259007,0.03291715495288372,0.03304482437670231,0.03317889105528593 +single,hard-sol-a-trap-04,1,0.025454387068748474,0.0253780884668231,0.02562911156564951,0.025792223401367664 +single,hard-sol-a-trap-13,1,0.025850597769021988,0.025490470230579376,0.025519649498164654,0.025620808824896812 +single,hard-sol-a-trap-15,1,0.03252379409968853,0.03231658972799778,0.03236565086990595,0.03212971054017544 +single,hard-sol-b-judge_hard-01,1,0.031981014646589756,0.03187849186360836,0.03182569518685341,0.03155826963484287 +single,hard-sol-b-judge_hard-02,1,0.02644006721675396,0.02668468188494444,0.026584510691463947,0.026384525001049042 +single,hard-sol-b-judge_hard-05,1,0.028231196105480194,0.028163609094917774,0.028525520116090775,0.02872964832931757 +single,hard-sol-b-judge_hard-08,1,0.026644441299140453,0.026962414383888245,0.026573524810373783,0.026577326469123363 +single,hard-sol-b-judge_hard-10,1,0.028205085545778275,0.02889356669038534,0.02756207063794136,0.029585218988358974 +single,hard-sol-b-judge_hard-14,1,0.027772091329097748,0.028051892295479774,0.027320955879986286,0.028694654814898968 +single,hard-sol-b-judge_hard-15,1,0.02664110530167818,0.02636315394192934,0.02652369625866413,0.027234056033194065 +single,hard-sol-b-judge_hard-18,1,0.024477320723235607,0.02382808458060026,0.02434912882745266,0.02495672833174467 +single,hard-sol-c-judge_hard-03,1,0.040552801452577114,0.040989965200424194,0.04189595114439726,0.04184152651578188 +single,hard-sol-c-judge_hard-05,1,0.03720196522772312,0.0387018658220768,0.03820089530199766,0.0378299867734313 +single,hard-sol-c-judge_hard-07,1,0.03605147358030081,0.036928389221429825,0.03605798538774252,0.03571134153753519 +single,hard-sol-c-judge_hard-08,1,0.03578027058392763,0.03331796079874039,0.034536113031208515,0.03319684974849224 +single,hard-sol-c-judge_hard-09,1,0.03959968499839306,0.037440359592437744,0.03772969637066126,0.03710372932255268 +single,hard-sol-c-judge_hard-10,1,0.040762814693152905,0.04035297594964504,0.040080199018120766,0.03984433505684137 +single,hard-sol-c-judge_hard-11,1,0.03815429285168648,0.03863897547125816,0.038152233697474,0.037671202793717384 +single,hard-sol-c-judge_hard-13,1,0.03701641317456961,0.03737915027886629,0.03697569761425257,0.03675248101353645 +single,hard-sol-c-judge_hard-15,1,0.03342884033918381,0.03331714868545532,0.03307904954999685,0.033179253339767456 +questions-4,original-policy-01-0-q4,4,0.08121675997972488,0.08077133819460869,0.06321162823587656,0.06507930066436529 +questions-4,original-policy-01-1-q4,4,0.08035885822027922,0.08028338756412268,0.06045289617031813,0.060953172855079174 +questions-4,original-policy-02-0-q4,4,0.08066147845238447,0.08092384785413742,0.05725836753845215,0.05729231610894203 +questions-4,original-policy-02-1-q4,4,0.08027491718530655,0.08027025312185287,0.0564371719956398,0.05635155364871025 +questions-4,original-policy-03-0-q4,4,0.08033824805170298,0.08040370978415012,0.05677344463765621,0.05692318081855774 +questions-4,original-policy-03-1-q4,4,0.0801288178190589,0.08024972211569548,0.05656735599040985,0.05681983381509781 +questions-4,original-policy-04-0-q4,4,0.0792480232194066,0.07938346546143293,0.05954006314277649,0.058688657358288765 +questions-4,original-policy-04-1-q4,4,0.08045385126024485,0.08079447504132986,0.058638128452003,0.05782437510788441 +questions-4,original-policy-05-0-q4,4,0.07963122148066759,0.07967977225780487,0.05990747641772032,0.05973853264003992 +questions-4,original-policy-05-1-q4,4,0.0802569342777133,0.08054290525615215,0.05743120424449444,0.057600575499236584 +questions-4,original-policy-06-0-q4,4,0.07976769097149372,0.07972896751016378,0.060274384915828705,0.06059269234538078 +questions-4,original-policy-06-1-q4,4,0.08052570093423128,0.08173689525574446,0.057524142786860466,0.058380674570798874 +questions-4,original-adequacy-01-0-q4,4,0.07914777100086212,0.07956595532596111,0.05471959616988897,0.055156827904284 +questions-4,original-adequacy-01-1-q4,4,0.0781913036480546,0.07826322875916958,0.05434275884181261,0.05419527646154165 +questions-4,original-adequacy-02-0-q4,4,0.0788645651191473,0.07889253925532103,0.05493144504725933,0.055393606424331665 +questions-4,original-adequacy-02-1-q4,4,0.07947022654116154,0.07942851353436708,0.05544907599687576,0.05547427572309971 +questions-4,original-adequacy-03-0-q4,4,0.07934652641415596,0.07940807938575745,0.05720856972038746,0.055893683806061745 +questions-4,original-adequacy-03-1-q4,4,0.07977482862770557,0.07952272985130548,0.056189331226050854,0.055645352229475975 +questions-4,original-adequacy-04-0-q4,4,0.07950487080961466,0.07942201942205429,0.0562599403783679,0.056472803466022015 +questions-4,original-adequacy-04-1-q4,4,0.07919340953230858,0.07915555126965046,0.055537400767207146,0.05508987884968519 +questions-4,original-adequacy-05-0-q4,4,0.07804075535386801,0.07804688345640898,0.05466490797698498,0.05443546827882528 +questions-4,original-adequacy-05-1-q4,4,0.07809603773057461,0.07806482445448637,0.054221078753471375,0.055259983986616135 +questions-4,original-adequacy-06-0-q4,4,0.08002882543951273,0.08000223059207201,0.05842426046729088,0.058377500623464584 +questions-4,original-adequacy-06-1-q4,4,0.0801912359893322,0.08014713600277901,0.05739328544586897,0.057219989597797394 +questions-4,easy-fact-00-q4,4,0.07925026677548885,0.07907824404537678,0.05438528675585985,0.0544439610093832 +questions-4,easy-fact-01-q4,4,0.07838347088545561,0.07826960645616055,0.055420104414224625,0.05412238836288452 +questions-4,easy-fact-02-q4,4,0.07955682463943958,0.07938146591186523,0.05494390707463026,0.05641800817102194 +questions-4,easy-fact-03-q4,4,0.0800684904679656,0.08006149530410767,0.056753477081656456,0.0563310319557786 +questions-4,easy-fact-04-q4,4,0.07833447400480509,0.07843679934740067,0.05441931542009115,0.053966631181538105 +questions-4,easy-fact-05-q4,4,0.07866242807358503,0.0787829514592886,0.05455311946570873,0.0548781082034111 +questions-4,easy-fact-06-q4,4,0.07931713946163654,0.07926564384251833,0.06197750475257635,0.06273819785565138 +questions-4,easy-fact-07-q4,4,0.07795443013310432,0.07792929373681545,0.05394017603248358,0.05396429356187582 +questions-4,easy-fact-08-q4,4,0.07779624499380589,0.0777341965585947,0.05558489914983511,0.05526716914027929 +questions-4,easy-fact-09-q4,4,0.07772835996001959,0.07778767216950655,0.05529818031936884,0.05482165236026049 +questions-4,easy-fact-10-q4,4,0.07856766227632761,0.07853275444358587,0.054916851222515106,0.054661222733557224 +questions-4,easy-fact-11-q4,4,0.07906865049153566,0.07909476384520531,0.056164815090596676,0.05467402748763561 +questions-4,hard-opus-a-probability-08-q4,4,0.15954467747360468,0.16167007759213448,0.1515036979690194,0.15133407432585955 +questions-4,hard-sol-a-adversarial-06-q4,4,0.11340197827666998,0.11018222291022539,0.1042754715308547,0.10880056489259005 +questions-4,hard-sol-a-adversarial-08-q4,4,0.10854428727179766,0.10908480081707239,0.09550650045275688,0.09719940461218357 +questions-4,hard-sol-a-trap-02-q4,4,0.13059690315276384,0.13163178134709597,0.12270038202404976,0.12002086453139782 +questions-4,hard-sol-a-trap-04-q4,4,0.10183833725750446,0.10134927090257406,0.0895042298361659,0.09015883784741163 +questions-4,hard-sol-a-trap-13-q4,4,0.1040317490696907,0.10315863322466612,0.09101295936852694,0.09177765063941479 +questions-4,hard-sol-a-trap-15-q4,4,0.13646090123802423,0.13204425666481256,0.12216869462281466,0.12436570320278406 +questions-4,hard-sol-b-judge_hard-01-q4,4,0.129834889434278,0.12748553603887558,0.1167262913659215,0.11798160430043936 +questions-4,hard-sol-b-judge_hard-02-q4,4,0.10448930785059929,0.10395772848278284,0.09223071672022343,0.09086407907307148 +questions-4,hard-sol-b-judge_hard-05-q4,4,0.11239486839622259,0.11311410553753376,0.10273012984544039,0.1008826196193695 +questions-4,hard-sol-b-judge_hard-08-q4,4,0.10841288324445486,0.10864491481333971,0.09251700341701508,0.09442747943103313 +questions-4,hard-sol-b-judge_hard-10-q4,4,0.11393897142261267,0.11464154627174139,0.10410753451287746,0.10386680532246828 +questions-4,hard-sol-b-judge_hard-14-q4,4,0.11327060963958502,0.1110995328053832,0.10640417318791151,0.10580927599221468 +questions-4,hard-sol-b-judge_hard-15-q4,4,0.10513919033110142,0.10643853526562452,0.09310258273035288,0.09522844385355711 +questions-4,hard-sol-b-judge_hard-18-q4,4,0.09400731883943081,0.09578841831535101,0.0793972471728921,0.08036240097135305 +questions-4,hard-sol-c-judge_hard-03-q4,4,0.16462202742695808,0.16286872047930956,0.15326612070202827,0.15335650742053986 +questions-4,hard-sol-c-judge_hard-05-q4,4,0.15439793188124895,0.15367948170751333,0.14211427047848701,0.14237547758966684 +questions-4,hard-sol-c-judge_hard-07-q4,4,0.1431005634367466,0.14196058362722397,0.1327682500705123,0.13217727094888687 +questions-4,hard-sol-c-judge_hard-08-q4,4,0.133301786147058,0.1329886382445693,0.12662024330347776,0.12686536740511656 +questions-4,hard-sol-c-judge_hard-09-q4,4,0.14702362194657326,0.14910012483596802,0.14203596115112305,0.1412455542013049 +questions-4,hard-sol-c-judge_hard-10-q4,4,0.16213070135563612,0.16103309486061335,0.15270814392715693,0.14912148378789425 +questions-4,hard-sol-c-judge_hard-11-q4,4,0.1486551184207201,0.14783655852079391,0.13909981027245522,0.14161898847669363 +questions-4,hard-sol-c-judge_hard-13-q4,4,0.1453861352056265,0.14731405023485422,0.13389879278838634,0.13447910640388727 +questions-4,hard-sol-c-judge_hard-15-q4,4,0.13022752664983273,0.13097235839813948,0.11986418161541224,0.11941787134855986 +questions-8,original-policy-01-0-q8,8,0.16078681591898203,0.1609477587044239,0.11501733679324389,0.11522812210023403 +questions-8,original-policy-01-1-q8,8,0.1599606405943632,0.16051357612013817,0.11690837237983942,0.1172329131513834 +questions-8,original-policy-02-0-q8,8,0.160559787414968,0.16061993036419153,0.11084581259638071,0.11060481239110231 +questions-8,original-policy-02-1-q8,8,0.1597910625860095,0.15997065883129835,0.10986006259918213,0.10980174038559198 +questions-8,original-policy-03-0-q8,8,0.15998746640980244,0.16032838076353073,0.11154317203909159,0.11142586078494787 +questions-8,original-policy-03-1-q8,8,0.16117733716964722,0.1611775914207101,0.10935582313686609,0.11007114220410585 +questions-8,original-policy-04-0-q8,8,0.16033499129116535,0.16036919597536325,0.11350606009364128,0.11372252367436886 +questions-8,original-policy-04-1-q8,8,0.160899942740798,0.1610167371109128,0.11706171091645956,0.11588556971400976 +questions-8,original-policy-05-0-q8,8,0.1597539521753788,0.16034169774502516,0.12045183219015598,0.12000620272010565 +questions-8,original-policy-05-1-q8,8,0.16117433924227953,0.1607238333672285,0.11343692801892757,0.11424795165657997 +questions-8,original-policy-06-0-q8,8,0.16079493053257465,0.1608268627896905,0.11732386332005262,0.11733606271445751 +questions-8,original-policy-06-1-q8,8,0.16074191126972437,0.16077054012566805,0.11267818789929152,0.11381306871771812 +questions-8,original-adequacy-01-0-q8,8,0.15870594419538975,0.15866100322455168,0.1067879656329751,0.10623183846473694 +questions-8,original-adequacy-01-1-q8,8,0.1562205646187067,0.1562682306393981,0.10601120255887508,0.10579135455191135 +questions-8,original-adequacy-02-0-q8,8,0.15728759299963713,0.15735474415123463,0.105084796436131,0.1065595829859376 +questions-8,original-adequacy-02-1-q8,8,0.15810890402644873,0.1579051474109292,0.1042238138616085,0.10505293030291796 +questions-8,original-adequacy-03-0-q8,8,0.15826172567903996,0.1585646914318204,0.10652301087975502,0.10615278501063585 +questions-8,original-adequacy-03-1-q8,8,0.15884616784751415,0.1587766008451581,0.1062777042388916,0.10819937661290169 +questions-8,original-adequacy-04-0-q8,8,0.1587467510253191,0.15881475806236267,0.1061132587492466,0.10852077044546604 +questions-8,original-adequacy-04-1-q8,8,0.15823656041175127,0.15838452149182558,0.10565267596393824,0.10837263613939285 +questions-8,original-adequacy-05-0-q8,8,0.1560490932315588,0.1558803990483284,0.1022888533771038,0.10460546892136335 +questions-8,original-adequacy-05-1-q8,8,0.15630687773227692,0.15603235643357038,0.10271507315337658,0.10341479163616896 +questions-8,original-adequacy-06-0-q8,8,0.16018626280128956,0.15998790226876736,0.11107498500496149,0.11142548732459545 +questions-8,original-adequacy-06-1-q8,8,0.15997202321887016,0.16024103481322527,0.11175504419952631,0.11116872821003199 +questions-8,easy-fact-00-q8,8,0.15751459263265133,0.1580251269042492,0.10498890187591314,0.10580422636121511 +questions-8,easy-fact-01-q8,8,0.156127592548728,0.15644090063869953,0.10577787552028894,0.10812567640095949 +questions-8,easy-fact-02-q8,8,0.15868798084557056,0.15865547116845846,0.10732526239007711,0.10818255878984928 +questions-8,easy-fact-03-q8,8,0.1645247731357813,0.16017436981201172,0.11111480463296175,0.11044393479824066 +questions-8,easy-fact-04-q8,8,0.15679533034563065,0.15635274909436703,0.10584187600761652,0.10683869756758213 +questions-8,easy-fact-05-q8,8,0.15861796960234642,0.15713563933968544,0.1041916674003005,0.10486364178359509 +questions-8,easy-fact-06-q8,8,0.15956519078463316,0.15940775629132986,0.11393707152456045,0.11446250323206186 +questions-8,easy-fact-07-q8,8,0.15637547802180052,0.15605615358799696,0.10599928349256516,0.10636500082910061 +questions-8,easy-fact-08-q8,8,0.15576528012752533,0.15583895612508059,0.10123438946902752,0.10207103565335274 +questions-8,easy-fact-09-q8,8,0.15608464181423187,0.15567784570157528,0.09992413315922022,0.10012270044535398 +questions-8,easy-fact-10-q8,8,0.15730905812233686,0.15738904476165771,0.10445038042962551,0.10417828988283873 +questions-8,easy-fact-11-q8,8,0.15803071670234203,0.15824270620942116,0.10466332361102104,0.10444131586700678 +questions-8,hard-opus-a-probability-08-q8,8,0.3227981962263584,0.32440505363047123,0.3059441354125738,0.30695053562521935 +questions-8,hard-sol-a-adversarial-06-q8,8,0.2305390015244484,0.23074593860656023,0.20472123380750418,0.2047848291695118 +questions-8,hard-sol-a-adversarial-08-q8,8,0.2185986414551735,0.22007994633167982,0.19076173286885023,0.1918493863195181 +questions-8,hard-sol-a-trap-02-q8,8,0.26008088514208794,0.25931433122605085,0.23842435237020254,0.24109140038490295 +questions-8,hard-sol-a-trap-04-q8,8,0.2017742618918419,0.20095088239759207,0.17280766367912292,0.1728870626538992 +questions-8,hard-sol-a-trap-13-q8,8,0.20535727497190237,0.2032159585505724,0.18157331179827452,0.1822934551164508 +questions-8,hard-sol-a-trap-15-q8,8,0.26925958693027496,0.2690282594412565,0.2451084228232503,0.2484713038429618 +questions-8,hard-sol-b-judge_hard-01-q8,8,0.2574078096076846,0.26066965982317924,0.23061509430408478,0.2297003148123622 +questions-8,hard-sol-b-judge_hard-02-q8,8,0.21303673647344112,0.2091723559424281,0.18816000316292048,0.18948931619524956 +questions-8,hard-sol-b-judge_hard-05-q8,8,0.22839554958045483,0.22839483246207237,0.2045498052611947,0.20244073122739792 +questions-8,hard-sol-b-judge_hard-08-q8,8,0.20912430807948112,0.2142850775271654,0.18658673483878374,0.18467264343053102 +questions-8,hard-sol-b-judge_hard-10-q8,8,0.2271367022767663,0.22553264908492565,0.2036987692117691,0.20484731812030077 +questions-8,hard-sol-b-judge_hard-14-q8,8,0.22732973098754883,0.22850064560770988,0.20405963715165854,0.20452484674751759 +questions-8,hard-sol-b-judge_hard-15-q8,8,0.20609948690980673,0.20389264449477196,0.18290913570672274,0.1833201888948679 +questions-8,hard-sol-b-judge_hard-18-q8,8,0.1878449134528637,0.18773649912327528,0.15711795445531607,0.15643949154764414 +questions-8,hard-sol-c-judge_hard-03-q8,8,0.3286048546433449,0.3302147137001157,0.3056042520329356,0.30241071432828903 +questions-8,hard-sol-c-judge_hard-05-q8,8,0.3102688295766711,0.31295446306467056,0.2797259949147701,0.28032094053924084 +questions-8,hard-sol-c-judge_hard-07-q8,8,0.2880890481173992,0.28815275616943836,0.25988804269582033,0.26010854821652174 +questions-8,hard-sol-c-judge_hard-08-q8,8,0.2727884389460087,0.2686002738773823,0.24796358402818441,0.24861462600529194 +questions-8,hard-sol-c-judge_hard-09-q8,8,0.29523557517677546,0.29829603992402554,0.27622724417597055,0.27648898866027594 +questions-8,hard-sol-c-judge_hard-10-q8,8,0.3169580725952983,0.3197693144902587,0.29039205703884363,0.2903343290090561 +questions-8,hard-sol-c-judge_hard-11-q8,8,0.29833577666431665,0.2989487200975418,0.276958592236042,0.2778787361457944 +questions-8,hard-sol-c-judge_hard-13-q8,8,0.29274905286729336,0.2900575948879123,0.2671976266428828,0.26990056317299604 +questions-8,hard-sol-c-judge_hard-15-q8,8,0.26207865308970213,0.2633320018649101,0.24056084360927343,0.23810052126646042 +mixed-types,choice-score-0,3,0.1367520811036229,0.1371784657239914,0.08931048680096865,0.08986332919448614 +mixed-types,choice-score-1,3,0.14127317164093256,0.14102720469236374,0.10503425356000662,0.1075960136950016 +mixed-types,choice-score-2,3,0.1616301704198122,0.16110578831285238,0.134822147898376,0.13644948601722717 +mixed-types,choice-score-3,3,0.18371393252164125,0.18325003795325756,0.1605533054098487,0.1624906938523054 +mixed-types,choice-score-4,3,0.2081585144624114,0.20976586360484362,0.18415218871086836,0.18341888301074505 +mixed-types,choice-score-5,3,0.24009016156196594,0.24020289443433285,0.20582004450261593,0.20914695225656033 +mixed-types,choice-score-6,3,0.2605044934898615,0.2590738916769624,0.2570664621889591,0.24886900559067726 +mixed-types,choice-score-7,3,0.28563778195530176,0.2838850812986493,0.26776767801493406,0.26990082021802664 +mixed-types,choice-score-8,3,0.3297129459679127,0.3358356477692723,0.2912302576005459,0.29130055475980043 +mixed-types,choice-score-9,3,0.35509475879371166,0.35277013201266527,0.33514653239399195,0.3341169795021415 +mixed-types,choice-score-10,3,0.39042334351688623,0.38130088802427053,0.3545829514041543,0.3550212234258652 +mixed-types,choice-score-11,3,0.40754649974405766,0.405721134506166,0.37989793717861176,0.37771117873489857 diff --git a/docs/architecture.md b/docs/architecture.md index 958ce00a..e4612f7a 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -9,7 +9,10 @@ design. Concrete input/output types follow each executor's supported layout. The [Rust frontend](../src/frontend/README.md) currently forwards HTTP requests to separately running workers. Cua-S1 and Open-Jev have native Rust/CUDA workers that share the [Qwen3.5/3.8 executor](../src/models/qwen3_5/native/), which accepts -one prompt per forward call. [Laya's native worker](../src/models/laya/README.md) +single prompts and bounded packed prefill. Cua-S1 uses single-prompt calls; +Open-Jev packs candidates within one request for input and gate/up GEMMs while +preserving per-sequence mixers and output/down GEMM shapes. +[Laya's native worker](../src/models/laya/README.md) uses a separate Hopper CUDA backend for one complete padded request. All three coordinate independent processors and executors through `prepare` → `execute` → `finish` and use the @@ -37,8 +40,10 @@ and candidate identity, usage, and response metadata outside the executor. | Open-Jev | Token-ID vectors grouped by question, then independent candidate, in request order. | One FP32 learned scalar per candidate in the same grouping. | Add the `noul` false logit of zero, calibrate across each complete question, and restore typed answers, usage, and metadata. | | Laya | One padded request: token IDs, true lengths, question types and ordered option markers; at most 16 questions, 512 tokens per row and 2048 markers. | Per-question FP32 option logits and two action logits copied back after GPU heads. | Calibrate and decode ordered `choice`, `score` and `noul` answers, usage and metadata. | -Qwen input collections are serial work, not GPU batches; Laya batches questions -within one request. Shared runtime +Cua-S1 input collections are serial work. Open-Jev's model-specific batch adapter +packs up to 16 independent candidates and 4096 tokens per group; longer prompts +execute alone. It restores question/candidate grouping before normalization. +Laya batches questions within one request. Shared runtime admission precedes blocking dispatch: Cua-S1 admits one question forward at a time; Open-Jev and Laya admit one complete request. Cua-S1's CPU letter projection stays outside admission; Open-Jev's scalar heads and Laya's GPU heads and diff --git a/recipe/open_jev/native.md b/recipe/open_jev/native.md index 7537d48f..245841e5 100644 --- a/recipe/open_jev/native.md +++ b/recipe/open_jev/native.md @@ -44,9 +44,8 @@ or an incomplete export. The saved limit defaults to 4096 tokens per candidate; The CUDA kernels require compute capability 8.0 or newer. The current build target below is Ada (`89`); pass your GPU's compute capability explicitly. -The CUDA shared library and both Rust workers must be rebuilt together because -the gated-attention entry point updates the library ABI to version 4 alongside -the shared CUDA Graph entry points. +The CUDA shared library and Rust workers must be rebuilt together for ABI +version 5, which includes the shared vision and CUDA Graph entry points. ```sh src/backends/cuda/qwen3_5/build.sh target/release 89 @@ -97,6 +96,12 @@ and CUDA kernel tests are opt-in; the latter require a GPU reservation: # Inside a GPU reservation, after building the library: CUA_S1_CUDA_LIB=$PWD/target/release/libqwen3_5_cuda.so \ cargo test --release --locked -p omni-qwen3-5-native --test kernels -- --ignored + +# Full-checkpoint packing, ordering and graph-shape checks, in the reservation: +OPEN_JEV_MODEL=$PWD/weights/open-jev-27b-merged \ +OPEN_JEV_CUDA_LIB=$PWD/target/release/libqwen3_5_cuda.so CUA_S1_GRAPH=1 \ + cargo test --release --locked -p omni-open-jev-native --test prefill_batch \ + -- --ignored --test-threads=1 ``` CPU golden fixtures come from Open-Jev's request compiler and response formatter @@ -115,14 +120,24 @@ RMSNorm keeps thread values in registers at widths 2560/5120. MLP SiLU uses 16-byte BF16 loads/stores when width, stride and pointers permit it, retaining both BF16 rounding points; other layouts use the scalar path. -This recipe leaves `CUA_S1_GRAPH` unset and runs one eager forward pass per -candidate. Set `CUA_S1_GRAPH=1` on the worker to enable CUDA Graph replay. The -shared backend retains at most 64 graphs, keyed by exact candidate token length; -growing the scratch buffer clears them. Capturing a new length first runs an -eager forward to initialize its plans, then captures and replays the forward. -This adds cost for new lengths, so graph mode remains opt-in. Warm replay is -validated on the 74-case H200 workload: its mean HTTP latency is 2.03% below -eager execution after all workload lengths are warmed. Tokenization, transfers +This recipe leaves `CUA_S1_GRAPH` unset and uses eager prefill. Candidates within +a request are packed in prepared order, up to 16 sequences and 4096 total tokens +per group; longer prompts execute alone without truncation. Input and gate/up +projections share GEMMs. Output/down projections preserve their per-prompt shapes +and reduction order, and each sequence retains independent attention, positions, +convolution and GDN state. Calibration still uses every candidate in its question. +The [H200 packing comparison](../../benchmarks/prefill_batching/README.md) +records latency, exact output checks and frozen controls. +Packing validation covers H200 (sm_90); other CUDA architectures remain unverified. + +Set `CUA_S1_GRAPH=1` on the worker to enable CUDA Graph replay. The +shared backend retains at most 64 graphs, keyed by ordered sequence token lengths; +growing the scratch buffer clears them. A new shape first runs an eager forward +to initialize its plans and captures the layer loop for later replay. +This adds cost for new lengths, so graph mode remains opt-in. The earlier +single-prompt graph comparison on the 74-case H200 workload measured mean HTTP +latency 2.03% below eager execution after all lengths were warmed. Combined +packing and graph performance remains unmeasured. Tokenization, transfers and the CPU scalar head remain outside the graph. Prefix sharing, GEMM autotuning, quantization and multimodal inference are not implemented. The [H200 validation](validation.md) reports full-checkpoint results for 74 diff --git a/src/backends/cuda/qwen3_5/README.md b/src/backends/cuda/qwen3_5/README.md index 27a9f9b3..e142176e 100644 --- a/src/backends/cuda/qwen3_5/README.md +++ b/src/backends/cuda/qwen3_5/README.md @@ -11,7 +11,12 @@ The norm, elementwise and q/k preparation kernels round to bfloat16 where Transf `cs1_attention_gated` fuses the sigmoid gate into the attention epilogue, preserving the BF16 rounding of both attention and sigmoid before multiplication. The native workers use this entry point; the separate operations remain available for kernel -comparisons. Rebuild the library and workers together for ABI version 4, which -includes the CUDA Graph entry points and gated attention. +comparisons. Rebuild the library and workers together for ABI version 5, which +includes the shared vision and CUDA Graph entry points alongside gated attention. Gated DeltaNet preparation stores converted TF32 operands in three-byte component planes, preserves the original four-term TF32 accumulation, and writes U/W fragments directly as bfloat16. Dynamic shared memory is 72 KiB per block. The [H200 comparison](../../../../benchmarks/gdn/README.md) records complete GDN call latency, numerical checks, and the small end-to-end change measured with the Open-Jev worker from PR #55. + +The shared Rust model can pack independent sequences for input and gate/up GEMMs. +Output/down GEMMs retain each prompt's original shape and reduction order; +attention, convolution and GDN calls remain sequence-local. The CUDA ABI is +unchanged. Open-Jev uses this path within requests; Cua-S1 keeps single-prompt calls. diff --git a/src/models/open_jev/README.md b/src/models/open_jev/README.md index a6b22822..01fe199a 100644 --- a/src/models/open_jev/README.md +++ b/src/models/open_jev/README.md @@ -24,7 +24,13 @@ by question and candidate, plus the response context for identity/order, calibration, usage, and metadata. The executor owns the trained head and returns one FP32 scalar per candidate. Response finishing adds the `noul` false baseline and applies calibrated normalization across each complete question. The worker -retains request-wide model locking and independent single-prompt execution, with -the scalar head on the CPU after CUDA prefill. The engine owns a +retains request-wide model locking and independent sequence semantics, with +the scalar head on the CPU after CUDA prefill. Its [batch adapter](native/src/batching.rs) +packs at most 16 candidates and 4096 tokens per group, in prepared order; +longer individual prompts execute alone without truncation. Input and gate/up +projections share packed GEMMs. Output/down projections retain their original +per-prompt GEMM shapes; attention, positions, convolution and GDN state reset +at each sequence boundary. Results are regrouped before question normalization. +The engine owns a [shared serial scheduler](../../runtime/README.md) that admits the complete request -before blocking dispatch. GPU batching and batch budgets remain planned. +before blocking dispatch. Cross-request batching and shared queue budgets remain planned. diff --git a/src/models/open_jev/native/Cargo.toml b/src/models/open_jev/native/Cargo.toml index 346440c6..52c15634 100644 --- a/src/models/open_jev/native/Cargo.toml +++ b/src/models/open_jev/native/Cargo.toml @@ -17,3 +17,7 @@ tokio = { version = "1.49.0", features = ["macros", "net", "rt-multi-thread", "s [[test]] name = "contract" path = "../../../../tests/open_jev/contract.rs" + +[[test]] +name = "prefill_batch" +path = "../../../../tests/open_jev/prefill_batch.rs" diff --git a/src/models/open_jev/native/src/batching.rs b/src/models/open_jev/native/src/batching.rs new file mode 100644 index 00000000..13f95c40 --- /dev/null +++ b/src/models/open_jev/native/src/batching.rs @@ -0,0 +1,29 @@ +//! Bounded packing of independent candidate prompts within one admitted request. + +use std::ops::Range; + +const MAX_SEQUENCES: usize = 16; +const MAX_TOKENS: usize = 4096; + +/// Keep candidate order, with long prompts executing alone at their original size. +pub(crate) fn ranges(inputs: &[&[u32]]) -> Vec> { + let mut batches = Vec::new(); + let mut start = 0; + let mut tokens = 0; + for (index, ids) in inputs.iter().enumerate() { + if index > start && (index - start == MAX_SEQUENCES || tokens + ids.len() > MAX_TOKENS) { + batches.push(start..index); + start = index; + tokens = 0; + } + tokens += ids.len(); + } + if start < inputs.len() { + batches.push(start..inputs.len()); + } + batches +} + +#[cfg(test)] +#[path = "../../../../../tests/open_jev/batching.rs"] +mod tests; diff --git a/src/models/open_jev/native/src/executor.rs b/src/models/open_jev/native/src/executor.rs index 9504c02e..cdccb641 100644 --- a/src/models/open_jev/native/src/executor.rs +++ b/src/models/open_jev/native/src/executor.rs @@ -8,6 +8,8 @@ use omni_qwen3_5_native::model::{Config, Model}; use omni_runtime::SerialScheduler; use serde_json::Value; +use crate::batching; + #[derive(Clone)] pub(crate) struct DecisionHead { weights: Vec, @@ -80,7 +82,7 @@ impl Executor { } /// Inputs and outputs are grouped by question, then candidate, in prepared order. - /// Admit one whole request; the model lock spans its independent candidate calls. + /// Admit one whole request and pack independent candidates for shared GEMMs. pub async fn execute( &self, scheduler: &SerialScheduler, @@ -93,14 +95,18 @@ impl Executor { let mut model = model .lock() .map_err(|_| anyhow::anyhow!("poisoned model"))?; - ids.iter() - .map(|candidates| { - candidates - .iter() - .map(|ids| head.score(model.forward(ids)?)) - .collect::>>() - }) - .collect::>>() + let inputs: Vec<&[u32]> = ids.iter().flatten().map(Vec::as_slice).collect(); + let mut scores = Vec::with_capacity(inputs.len()); + for range in batching::ranges(&inputs) { + for last in model.forward_batch(&inputs[range])? { + scores.push(head.score(last)?); + } + } + let mut scores = scores.into_iter(); + Ok(ids + .iter() + .map(|candidates| scores.by_ref().take(candidates.len()).collect()) + .collect()) }) .await } diff --git a/src/models/open_jev/native/src/lib.rs b/src/models/open_jev/native/src/lib.rs index 857f53c1..268ac289 100644 --- a/src/models/open_jev/native/src/lib.rs +++ b/src/models/open_jev/native/src/lib.rs @@ -1,4 +1,5 @@ //! Open-Jev-27B-v1.1 request compilation, candidate scoring and typed responses. +mod batching; pub mod contract; pub mod engine; pub mod executor; diff --git a/src/models/qwen3_5/native/src/model.rs b/src/models/qwen3_5/native/src/model.rs index 87952e6a..6d37cd2f 100644 --- a/src/models/qwen3_5/native/src/model.rs +++ b/src/models/qwen3_5/native/src/model.rs @@ -561,11 +561,11 @@ pub struct Model { layers: Vec, stream: Stream, gemm: *mut c_void, - /// Buffers for the longest prompt so far; grows as needed. + /// Buffers for the largest packed token count so far; grows as needed. scratch: Option, - /// Opt-in replay with at most eight exact-length captures. + /// Opt-in replay with at most 64 captures keyed by ordered sequence lengths. graph_enabled: bool, - graphs: VecDeque<(usize, cuda::Graph)>, + graphs: VecDeque<(Vec, cuda::Graph)>, } // SAFETY: the raw pointers are device addresses and a cuBLASLt handle owned by the @@ -681,6 +681,24 @@ impl Model { ) } + /// Preserve each prompt's output/down GEMM shape and split-K reduction order. + fn gemm_sequences( + &self, + s: &Scratch, + x: usize, + w: &Tensor, + y: usize, + lengths: &[usize], + ) -> Result<()> { + let (n, k) = (w.shape[0], w.shape[1]); + let mut offset = 0; + for &length in lengths { + self.gemm(s, x + offset * k * BF16, w, y + offset * n * BF16, length)?; + offset += length; + } + Ok(()) + } + /// Finish queued work before releasing external execution admission. pub fn synchronize(&self) -> Result<()> { cuda::set_device(0)?; @@ -703,36 +721,55 @@ impl Model { /// The final-norm hidden state at the last position, as float32. pub fn forward(&mut self, ids: &[u32]) -> Result> { - let t = ids.len(); + Ok(self.forward_batch(&[ids])?.pop().unwrap()) + } + + /// Pack independent text prompts for input and gate/up GEMMs. + /// Mixers reset at each boundary; final-position hidden states retain input order. + pub fn forward_batch(&mut self, inputs: &[&[u32]]) -> Result>> { + ensure!(!inputs.is_empty(), "empty batch"); ensure!( - t > 0 && t <= self.cfg.max_positions, + inputs + .iter() + .all(|ids| !ids.is_empty() && ids.len() <= self.cfg.max_positions), "empty or oversized prompt" ); ensure!( - ids.iter().all(|&i| (i as usize) < self.embed.shape[0]), + inputs + .iter() + .flat_map(|ids| ids.iter()) + .all(|&id| (id as usize) < self.embed.shape[0]), "token id outside the vocabulary" ); + let lengths: Vec = inputs.iter().map(|ids| ids.len()).collect(); + let t = lengths + .iter() + .try_fold(0usize, |total, &length| total.checked_add(length)) + .context("packed token count overflow")?; + ensure!( + t <= i32::MAX as usize, + "packed token count exceeds the CUDA layout" + ); self.prepare_scratch(t)?; let s = self.scratch.as_ref().unwrap(); - self.embed_tokens(s, ids)?; + let ids: Vec = inputs.iter().flat_map(|ids| ids.iter().copied()).collect(); + self.embed_tokens(s, &ids)?; if self.graph_enabled { - if let Some((_, graph)) = self.graphs.iter().find(|(length, _)| *length == t) { + if let Some((_, graph)) = self.graphs.iter().find(|(shape, _)| *shape == lengths) { graph.launch(self.stream)?; } else { - // Warm GEMM plans and keep this eager result for the cache miss. - // run() advances s.res in place and no longer embeds tokens, so - // launching the new graph here would advance the residual twice. - self.run(s, t, false)?; + // Warm plans and keep the eager result: run() advances the residual + // in place, so replaying on this cache miss would advance it twice. + self.run(s, &lengths, false)?; cuda::synchronize(self.stream)?; - match cuda::Graph::capture(self.stream, || self.run(s, t, false)) { + match cuda::Graph::capture(self.stream, || self.run(s, &lengths, false)) { Ok(graph) => { if self.graphs.len() == 64 { self.graphs.pop_front(); } - self.graphs.push_back((t, graph)); + self.graphs.push_back((lengths.clone(), graph)); } Err(error) => { - // Capture records without executing: the eager result is valid. eprintln!("CUDA Graph capture failed; using eager execution: {error:#}"); self.graph_enabled = false; self.graphs.clear(); @@ -740,9 +777,16 @@ impl Model { } } } else { - self.run(s, t, false)?; + self.run(s, &lengths, false)?; } - self.last_hidden(s, t) + let mut end = 0; + lengths + .iter() + .map(|&length| { + end += length; + self.last_hidden(s, end) + }) + .collect() } /// Prefill one unpadded prompt with already-adapted BF16 image embeddings and @@ -789,7 +833,7 @@ impl Model { } begin = end; } - self.run(s, t, true)?; + self.run(s, &[t], true)?; self.last_hidden(s, t) } @@ -845,10 +889,19 @@ impl Model { .collect()) } - /// Queue language layers over prepared embeddings in `s.res`, with rotary - /// tables in immutable text buffers or separate explicit-position buffers. + /// Queue language layers over packed embeddings, resetting sequence positions. /// Final-norm hidden states end up in `s.x`. - fn run(&self, s: &Scratch, t: usize, custom_positions: bool) -> Result<()> { + fn run(&self, s: &Scratch, lengths: &[usize], custom_positions: bool) -> Result<()> { + let t: usize = lengths.iter().sum(); + let mut start = 0; + let sequences: Vec<(usize, i32)> = lengths + .iter() + .map(|&length| { + let offset = start; + start += length; + (offset, length as i32) + }) + .collect(); let cfg = &self.cfg; let st = self.stream; let (ti, hi, eps) = (t as i32, cfg.hidden as i32, cfg.eps); @@ -861,7 +914,7 @@ impl Model { } else { (s.cos, s.sin) }; - // SAFETY (every kernel call below): pointers are weights in the arena or + // SAFETY (every kernel call below): the pointers are weights in the arena or // scratch buffers laid out for at least t tokens with the widths used here. unsafe { check( @@ -886,21 +939,23 @@ impl Model { let b = z + vd * BF16; let a = b + hv * BF16; unsafe { - check( - (cuda::api().cs1_gdn_conv)( - p(s.gdn_in), - ld, - la.conv.ptr, - p(s.lq), - p(s.lk), - p(s.lv), - ti, - kd as i32, - vd as i32, - st, - ), - "gdn conv", - )?; + for &(offset, length) in &sequences { + check( + (cuda::api().cs1_gdn_conv)( + p(s.gdn_in + offset * w.gdn_in * BF16), + ld, + la.conv.ptr, + p(s.lq + offset * kd * BF16), + p(s.lk + offset * kd * BF16), + p(s.lv + offset * vd * BF16), + length, + kd as i32, + vd as i32, + st, + ), + "gdn conv", + )?; + } check( (cuda::api().cs1_gdn_gates)( p(b), @@ -916,23 +971,25 @@ impl Model { ), "gdn gates", )?; - check( - (cuda::api().cs1_gdn_prefill)( - p(s.lq), - p(s.lk), - p(s.lv), - p(s.g).cast(), - p(s.beta), - p(s.lo), - p(s.workspace).cast(), - ti, - hv as i32, - cfg.lin_k_heads as i32, - (cfg.lin_k_dim as f32).powf(-0.5), - st, - ), - "gdn prefill", - )?; + for &(offset, length) in &sequences { + check( + (cuda::api().cs1_gdn_prefill)( + p(s.lq + offset * kd * BF16), + p(s.lk + offset * kd * BF16), + p(s.lv + offset * vd * BF16), + p(s.g + offset * hv * F32).cast(), + p(s.beta + offset * hv * BF16), + p(s.lo + offset * vd * BF16), + p(s.workspace).cast(), + length, + hv as i32, + cfg.lin_k_heads as i32, + (cfg.lin_k_dim as f32).powf(-0.5), + st, + ), + "gdn prefill", + )?; + } check( (cuda::api().cs1_gated_rms_norm)( p(s.lo), @@ -949,7 +1006,7 @@ impl Model { "gated norm", )?; } - self.gemm(s, s.ln, &la.out, s.delta, t)?; + self.gemm_sequences(s, s.ln, &la.out, s.delta, lengths)?; } Mixer::Full(fa) => { self.gemm(s, s.x, &fa.qkv, s.attn_in, t)?; @@ -957,47 +1014,49 @@ impl Model { let k = s.attn_in + w.attn_q * BF16; let v = k + cfg.kv_heads * cfg.head_dim * BF16; unsafe { - check( - (cuda::api().cs1_attn_prep)( - p(s.attn_in), - p(k), - ld, - fa.q_norm.ptr, - fa.k_norm.ptr, - p(cos), - p(sin), - p(s.aq), - p(s.agate), - p(s.ak), - ti, - hq, - hk, - hd, - cfg.rotary_half as i32, - eps, - st, - ), - "attention prep", - )?; - check( - (cuda::api().cs1_attention_gated)( - p(s.aq), - p(s.ak), - p(v), - ld, - p(s.agate), - p(s.ao), - ti, - hq, - hk, - hd, - (cfg.head_dim as f32).powf(-0.5), - st, - ), - "gated attention", - )?; + for &(offset, length) in &sequences { + check( + (cuda::api().cs1_attn_prep)( + p(s.attn_in + offset * w.attn_in * BF16), + p(k + offset * w.attn_in * BF16), + ld, + fa.q_norm.ptr, + fa.k_norm.ptr, + p(cos), + p(sin), + p(s.aq + offset * cfg.heads * cfg.head_dim * BF16), + p(s.agate + offset * cfg.heads * cfg.head_dim * BF16), + p(s.ak + offset * cfg.kv_heads * cfg.head_dim * BF16), + length, + hq, + hk, + hd, + cfg.rotary_half as i32, + eps, + st, + ), + "attention prep", + )?; + check( + (cuda::api().cs1_attention_gated)( + p(s.aq + offset * cfg.heads * cfg.head_dim * BF16), + p(s.ak + offset * cfg.kv_heads * cfg.head_dim * BF16), + p(v + offset * w.attn_in * BF16), + ld, + p(s.agate + offset * cfg.heads * cfg.head_dim * BF16), + p(s.ao + offset * cfg.heads * cfg.head_dim * BF16), + length, + hq, + hk, + hd, + (cfg.head_dim as f32).powf(-0.5), + st, + ), + "gated attention", + )?; + } } - self.gemm(s, s.ao, &fa.o, s.delta, t)?; + self.gemm_sequences(s, s.ao, &fa.o, s.delta, lengths)?; } } unsafe { @@ -1029,7 +1088,7 @@ impl Model { "silu mul", )?; } - self.gemm(s, s.act, &layer.down, s.delta, t)?; + self.gemm_sequences(s, s.act, &layer.down, s.delta, lengths)?; let next = self .layers .get(i + 1) diff --git a/tests/open_jev/batching.rs b/tests/open_jev/batching.rs new file mode 100644 index 00000000..b9632330 --- /dev/null +++ b/tests/open_jev/batching.rs @@ -0,0 +1,17 @@ +use super::ranges; + +#[test] +fn preserves_all_candidates_across_token_and_sequence_limits() { + let prompts: Vec> = [2048, 2048, 1, 5000, 3] + .into_iter() + .chain(std::iter::repeat_n(4, 17)) + .map(|length| vec![0; length]) + .collect(); + let inputs: Vec<&[u32]> = prompts.iter().map(Vec::as_slice).collect(); + let batches = ranges(&inputs); + assert_eq!(batches, vec![0..2, 2..3, 3..4, 4..20, 20..22]); + assert_eq!( + batches.into_iter().flatten().collect::>(), + (0..inputs.len()).collect::>() + ); +} diff --git a/tests/open_jev/prefill_batch.rs b/tests/open_jev/prefill_batch.rs new file mode 100644 index 00000000..4fbfad15 --- /dev/null +++ b/tests/open_jev/prefill_batch.rs @@ -0,0 +1,98 @@ +//! Full-checkpoint packing, sequence isolation and graph-shape regression check. +//! Run inside a GPU reservation with OPEN_JEV_MODEL and OPEN_JEV_CUDA_LIB set. + +use std::path::PathBuf; + +use omni_open_jev_native::engine::Engine; +use serde_json::json; + +#[tokio::test(flavor = "current_thread")] +#[ignore = "needs the pinned Open-Jev export, CUDA library and a GPU reservation"] +async fn packed_candidates_preserve_isolation_order_and_graph_shapes() { + let model = PathBuf::from(std::env::var_os("OPEN_JEV_MODEL").unwrap()); + let library = PathBuf::from(std::env::var_os("OPEN_JEV_CUDA_LIB").unwrap()); + let engine = Engine::load(&model, &library).await.unwrap(); + let request = json!({ + "state": "Refunds require a receipt. This customer has no receipt.", + "questions": { + "short": {"type": "noul", "instructions": "Is a refund permitted?"}, + "long": {"type": "noul", "instructions": format!("{}Is a refund permitted?", "Use only the stated policy. ".repeat(20))}, + "other": {"type": "noul", "instructions": "Does the customer have the required receipt?"} + } + }); + let prepared = engine + .processor + .prepare(&serde_json::to_vec(&request).unwrap()) + .unwrap(); + let inputs = prepared.inputs; + assert_ne!(inputs[0][0].len(), inputs[1][0].len()); + let mut independent = Vec::new(); + for candidate in &inputs { + independent.push( + engine + .executor + .execute(&engine.scheduler, vec![candidate.clone()]) + .await + .unwrap()[0][0], + ); + } + for order in [ + vec![0, 1, 2], + vec![1, 0, 2], + vec![2, 1, 0], + vec![0, 1, 2], + (0..17).map(|index| index % 3).collect(), + vec![0, 1, 2], + vec![0, 1, 2], + ] { + let grouped = order.iter().map(|&index| inputs[index].clone()).collect(); + let packed = engine + .executor + .execute(&engine.scheduler, grouped) + .await + .unwrap(); + assert_eq!(packed.len(), order.len()); + for (row, index) in packed.iter().zip(order) { + assert_eq!(row.len(), 1); + // Probability gates in the HTTP benchmark are stricter and cover + // complete typed outputs. This logit check detects sequence leakage. + assert!( + (row[0] - independent[index]).abs() <= 0.01, + "candidate {index}: packed={} independent={}", + row[0], + independent[index] + ); + } + } + // A later singleton must not retain a packed sequence's recurrent state. + let again = engine + .executor + .execute(&engine.scheduler, vec![inputs[0].clone()]) + .await + .unwrap(); + assert_eq!(again[0][0], independent[0]); + + // Reuse both singleton and packed graph shapes with changed token IDs. + // The new embeddings must replace the residual left by the previous replay. + let mut changed = inputs.clone(); + changed[0][0].fill(42); + let expected = engine + .executor + .execute(&engine.scheduler, vec![changed[0].clone()]) + .await + .unwrap()[0][0]; + assert_ne!(expected, independent[0]); + let packed = engine + .executor + .execute(&engine.scheduler, changed) + .await + .unwrap(); + assert_eq!(packed.len(), 3); + for (row, expected) in packed + .iter() + .zip([expected, independent[1], independent[2]]) + { + assert_eq!(row.len(), 1); + assert!((row[0] - expected).abs() <= 0.01); + } +}