Conversation
Map E4M3 and BOOL through the existing ATen adaptor and preserve the caller's CUDA device across NCCL communicator destruction. Reuse the existing paged Prefill warp kernel for NVIDIA head size 256, without changing other vendors' default dispatch. Extend existing multi-page/long-context coverage and add finite FP8/mask cast checks. Validation: fresh SM86 build, 88 paged Prefill cases, 2 cast tests, and TP2 communicator teardown from both caller devices. Closes InfiniTensor#1565
Consolidate the allocator ownership, reduction/scalar-power recording and MetaX concatenation fixes from InfiniTensor#1560 into the runtime support PR. Preserve the original implementations and focused regressions without adding a separate compiler or Prefill graph path.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem and result
Closes #1565. Target:
InfiniLM-v0.2.9c.Qwen greedy MTP needs E4M3/BOOL conversions, correct TP communicator teardown, head-256 paged Prefill, and reliable graph storage/replay. This PR supplies the complete shared-runtime prerequisites for InfiniTensor/InfiniLM#584. Its graph fixes are also used by Mamba-2 in InfiniTensor/InfiniLM#575.
Implementation
ncclCommDestroy.The implementation is organized into two reviewable commits: operator/communication support, then graph prerequisites. The latter consolidates #1560 without changing its code or tests. No Mamba scan, Prefill compiler, model weights or experiment artifacts are included. Native W8A8 is not introduced.
Validation
Current-head local verification (2026-09-23): rebuilt and installed
a3ac4df4for NVIDIA SM86 with graph support on the A6000 server. 13 graph-replay/FP8/BOOL checks and 88 paged Prefill cases passed. Paired with InfiniLM #58431ef4290, the small hybrid fixture passes TP1 (66 passed, one TP2-only skip) and TP2 (67 passed); LM totals include CPU checks. A stale September 19 installed library initially caused a graph allocator abort; the same reproducer passes after installing this PR's current library, without source/test changes. Current logs, hashes, conditions and commands. These shared-GPU correctness checks do not rerun the real 27B model or throughput benchmarks.Earlier platform evidence:
Consolidation audit and new focused logs; A6000 evidence; RTX 5090 evidence. Evidence lives on independent fork documentation branches.
The 5090 runtime's graph production changes are now inside this PR; SM120 build-option enumeration remains a documented validation overlay. Builds/tests above retain their exact recorded revisions. The earlier consolidation checks reused the matching prebuilt runtime; the September 23 checks above rebuilt the current head. Project formatting (clang-format 21.1.8/Ruff 0.15.20) and whitespace checks pass.
Performance and limitations
On TP2 RTX 5090, real 27B FP8 weights/BF16 compute, batch=1 greedy, 80×64-token pages, 63/127/1023-input and 64-output tokens, three warm repeats: ordinary Decode graph gives 63.21/62.43/52.16 tok/s; K2 eager MTP gives 107.25/116.50/86.79. These are whole-stack MTP gains, not isolated Core-patch speedups. Sampled K2 device peaks are 24394/24398 MiB (100 ms sampling).
Resolved integration issue: the previously reported 5090 cancel/re-admit mismatch at token 21 was fixed in InfiniTensor/InfiniLM#584, production commit
78d19f74. Aligning the per-token GDN gate projection arithmetic between batched Decode and checkpointed verification resolved the reproduced divergence. The original controlled lifecycle test now passes for K1/K2/K4, including cancellation, re-admission and complete KV/state reclamation; the same-state TP2 probe matches all 384 Conv/GDN tensors and both hidden vectors exactly. No comparison was relaxed, no target replay was added, and this Core implementation is unchanged. Fix, reproducer and verification.vLLM K2 failed its own ordinary-output comparison in the archived experiment; no validated cross-framework MTP speed claim is made.
Review and CI
Head
a3ac4df4: fork Ruff passed; Linux/Windows build matrix passed all four Ubuntu/Windows build and CPU-test jobs. The existing workflow labels jobs debug/release but does not passmatrix.typeto the build command; these results do not certify distinct build modes or GPU execution. Previous original-head build matrices passed. Upstream external-PR jobs need maintainer approval; no upstream success is implied.Please review allocator/graph ownership separately from numerical dispatch. @PanZezhong1725 and @spike-zhu were identified for review; formal assignment previously failed for lack of upstream permissions. Keep approval and required CI as merge gates. The MetaX code is supported by the linked C500 evidence, not a new vendor-wide validation claim.