Skip to content

[None][refactor] BREAKING: Remove the Python backend of KVCacheManagerV2 - #19154

Open
lowsfer wants to merge 2 commits into
NVIDIA:mainfrom
lowsfer:kvcm2-remove-python-backend
Open

lowsfer wants to merge 2 commits into
NVIDIA:mainfrom
lowsfer:kvcm2-remove-python-backend

Conversation

@lowsfer

@lowsfer lowsfer commented Sep 14, 2026 •

Copy link
Copy Markdown
Member

Summary

KVCacheManagerV2 shipped two implementations of the same subsystem behind
TLLM_KV_CACHE_MANAGER_V2_BACKEND. The C++ port is the default, is what CI exercises, and
is the only backend newer features support, so the Python implementation was carrying
duplicate block-key hashing, eviction and stats logic that had to stay bit-identical to
C++, plus a mypyc + rawref build pipeline that existed only to make it fast enough to
matter.

tensorrt_llm/runtime/kv_cache_manager_v2/ is now a re-export shim over the nanobind
module plus the backend-agnostic _introspection dispatcher: 36 tracked files down to 4,
net −13k lines.

Commits

  1. Remove the Python backend — the deletion, consumer migration and build cleanup.
  2. Lock SSM life cycles into KVCM2 page-movement statistics — new C++ coverage for
    behaviour that only the deleted Python tests guarded.

New bindings

Only one consumer needed native support. The KV-aware router's V2 hashing is expressed with
the existing sequence_to_blockchain_keys (every caller chains from a reuse-scope root, and
its first yielded pair is that root), so v2_sha256_block_hasher is gone and no hashing
binding was added
. Equivalence was verified against a reference implementation of the old
per-chunk chaining across 3 block sizes x 2 algorithms x salted/unsalted -- byte-identical --
and against golden digests captured from the Python hasher before deleting it.

The native disaggregated bounce buffer did need something: PooledPhysMemAllocator and
VirtMem wrap the existing cudaVirtMem.{h,cpp}. They register on the _introspection
submodule
, not the package surface, because they carry no stability promise -- the package
namespace is the stable surface. Three members total (device_id, address, destroy); the
package's exported API shrinks from 74 to 72 symbols.

BREAKING

  • TLLM_KV_CACHE_MANAGER_V2_BACKEND=python no longer exists.
  • build_wheel.py --mypyc and TRTLLM_ENABLE_MYPYC are gone, along with setup_mypyc.py,
    the rawref C extension and the setup.py packaging surgery they required.
  • Streaming KV events (kv_cache_config.kv_events_config) are dropped. The sink is
    duck-typed Python and the C++ radix tree calls its sink natively, so there is no live
    path. validate_streaming_support now rejects the config and points at the buffered path
    via event_buffer_max_size. The interface is kept as a stub and its tests are skipped
    rather than deleted, so a native sink can restore it later.

Test coverage

Deleting the Python implementation orphaned test_kv_cache_stats_life_cycles.py, which
drove the Python page-movement recorders through a duck-typed stand-in. That behaviour —
SSM/recurrent life cycles appearing in iteration stats, with global cache-hit counters
staying attention-only — was fixed in both backends by #17447, but only ever tested in
Python, and no C++ test builds an SSM life cycle at all.

The second commit closes that with a hybrid attention + SSM fixture (the first in the C++
suite) and three cases: offload/onboard per life cycle, the attention-only alloc-counter
guard, and host drops. Each was confirmed to fail when the life-cycle filters are restored
in KvCache::_recordMigratedSlots / _recordDroppedPages
— a regression lock that cannot
detect the regression is worthless.

Host-drop coverage is new for attention too; nothing asserted it against a live manager
before.

Verification

  • kvCacheManagerV2StatsTest: 13/13 (10 pre-existing + 3 new)
  • tests/unittest/kv_cache_manager_v2_tests/: 255 passed, 20 skipped
  • tests/unittest/_torch/executor/kv_cache/: 788 passed
  • executor/test_stats_serializer.py, disaggregated/test_router.py: green
  • Repo-wide grep for any deleted submodule: clean

All run on a B200. Note the local dev box is an H100 while cpp/build was configured
CUDA_ARCHITECTURES=100-real, which aborts on kernel launch — an artefact of the build
config, not a code defect, and it reproduces identically on unmodified main.

Dev Engineer Review

  • Removes the Python KVCacheManagerV2 backend, backend selection, rawref, mypyc build support, and Python implementation modules.
  • Loads the C++ implementation directly and exposes native PooledPhysMemAllocator and VirtMem.
  • Rejects streaming KV events and directs users to buffered events through event_buffer_max_size.
  • Removes private APIs and changes router hashing to use sequence_to_blockchain_keys. Verify downstream hash compatibility.
  • Adds hybrid attention and SSM life-cycle statistics coverage.
  • Review finding counts are unavailable from the supplied evidence.

QA Engineer Review

  • Updates KV cache manager, event, hashing, statistics, CUDA, attention, executor, and modeling tests.
  • Removes backend-specific skips, Python-only allocator tests, obsolete statistics fixtures, and streaming implementation coverage.
  • Adds coverage for hybrid attention and SSM page movement, allocation accounting, scoped blockchain keys, virtual-memory lifetime, and native error handling.
  • Streaming tests verify rejection and buffered-event guidance.
  • Reported verification includes 13/13 C++ statistics tests, 255 passed and 20 skipped KV cache manager tests, and 788 executor tests. A later merge pipeline failed, so overall CI status needs follow-up.
  • No test-list files changed. KV cache manager tests are registered in l0_h100.yml, l0_b200.yml, l0_cpu.yml, and l0_a10.yml. Coverage verdict: needs follow-up.

Per-File QA Perspective

  • .gitignore: Removes Python and mypyc artifact exclusions. Verify obsolete artifacts cannot enter source or packaging workflows.
  • C++ guide and exception files: Update native test guidance and Python exception documentation. Verify configuration errors map to Python AssertionError.
  • Nanobind binding: Adds PooledPhysMemAllocator and VirtMem. Verify construction, lifetime retention, address access, device ID access, and destruction.
  • C++ statistics test and test utility: Add hybrid attention and SSM tiered-cache coverage. Verify execution in the C++ test target.
  • Documentation and example files: Document unsupported streaming events and remove the Python-backend restriction from NVFP4 cold-page setup. Verify documentation matches runtime behavior.
  • Build and packaging files: Remove mypyc, rawref, and Python KV cache manager packaging paths. Verify wheel contents and obsolete build options.
  • Attention, disaggregation, and executor source files: Move imports to public or _introspection exports and remove backend-specific validation. Verify imports, disaggregation allocation, and NVFP4 validation.
  • Streaming event source files: Replace streaming behavior with unsupported-operation stubs. Verify rejection, error precedence, and buffered-event behavior.
  • Runtime package files: Make C++ bindings unconditional and remove Python implementation modules, private APIs, rawref, and type stubs. Verify public API compatibility and native binding availability.
  • Router files: Replace the standalone V2 hasher with chained keys. Verify V2 and V2-SHA256-64 compatibility and tail-rewrite behavior.
  • Performance YAML: Removes the backend environment override. Verify workers use the native default.
  • Test support files: Add shared CUDA utilities and update imports and temporary path handling. Verify CUDA error mapping, stream cleanup, and event lifetime.
  • KV cache manager tests: Remove backend gating and Python-only coverage. Verify native APIs, resizing, codecs, transfers, statistics, hashing, and regression paths.
  • Streaming tests: Skip unsupported streaming construction and verify rejection guidance. Verify buffered-event regressions remain covered.
  • Integration test: Uses direct CUDA initialization. Verify setup remains reliable in integration environments.
  • Test-list registration: No list files changed. Existing KV cache manager coverage is registered in the identified test-db suites.

@lowsfer
lowsfer force-pushed the kvcm2-remove-python-backend branch 5 times, most recently from 40a5e12 to 6493b80 Compare September 17, 2026 03:46
@lowsfer
lowsfer marked this pull request as ready for review September 17, 2026 04:07
@lowsfer
lowsfer requested review from a team as code owners September 17, 2026 04:07
@lowsfer

lowsfer commented Sep 17, 2026

Copy link
Copy Markdown
Member Author

/bot run --disable-fail-fast

Comment thread setup.py
with open("README.md", "r", encoding="utf-8") as fh:
long_description = fh.read()

# We use find_packages with a custom exclude filter to handle the mypyc compiled modules.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No — those should stay. Different tools with confusingly similar names:

  • mypyc is the ahead-of-time compiler that built the pure-Python KVCacheManagerV2 into a .so. That is what this PR removes (--mypyc, TRTLLM_ENABLE_MYPYC, setup_mypyc.py, and the setup.py packaging surgery they needed). Zero mypyc references remain in tracked files.
  • mypy is the static type checker. Both links you posted run it via scripts/run_mypy.sh — Build.groovy's typeCheck stage and the type-check pre-commit hook — and neither has anything to do with mypyc.

The only pyproject.toml mypy change here is dropping the [[tool.mypy.overrides]] block for tensorrt_llm.runtime.kv_cache_manager_v2.*, whose disallow_any_generics = false existed to accommodate the deleted Python implementation.

If anything the type check now covers slightly more: the PR restores runtime/kv_cache_manager_v2/__init__.pyi and fills in 12 exported names the stub was missing, so mypy --strict has more to read, not less.

if config.algorithm == "quantization_for_cold_page":
from tensorrt_llm.runtime.kv_cache_manager_v2 import _BACKEND

if _BACKEND == "python":

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removing _BACKEND and this admission check also requires migrating test_quantization_for_cold_page.py. Four tests still start with monkeypatch.setattr(runtime_v2_mod, "_BACKEND", ...), which now raises AttributeError before their assertions. The first also expects the deleted Python-backend rejection, so raising=False alone would still fail. Please remove those obsolete patches/assertions while retaining the SM100, speculative-mode, and estimation checks.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Valid — confirmed by running them, all four failed as you described. Fixed.

One correction for anyone following along: the file is tests/unittest/_torch/kv_cache_compression/test_quantization_for_cold_page.py (not _torch/executor/kv_cache/). Everything else matched, including the detail that raising=False alone would not have been enough for the first test, since it asserts the deleted Python-backend rejection.

Changes: dropped that rejection block, removed the six now-meaningless monkeypatch.setattr(runtime_v2_mod, "_BACKEND", "cpp") pins, and cleaned up what they orphaned — the runtime_v2_mod import and the monkeypatch fixture parameter on the two tests that no longer use it. The SM100, speculative-mode and estimation checks are all retained as you asked.

47 passed in that file (was 4 failing).

This one is squarely my miss: my post-removal sweep greps module paths rather than symbol names, so a _BACKEND attribute reference in a directory I was not running never surfaced. Thanks for catching it.

Comment thread .gitignore
@lowsfer
lowsfer force-pushed the kvcm2-remove-python-backend branch 2 times, most recently from ce691ea to aad3c39 Compare September 27, 2026 09:06
@lowsfer

lowsfer commented Sep 27, 2026

Copy link
Copy Markdown
Member Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #75485 [ run ] triggered by Bot. Commit: aad3c39 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #75485 [ run ] completed with state SUCCESS. Commit: aad3c39
/LLM/main/L0_MergeRequest_PR pipeline #62218 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@lowsfer
lowsfer force-pushed the kvcm2-remove-python-backend branch from aad3c39 to c26cbfd Compare September 28, 2026 12:48
@lowsfer

lowsfer commented Sep 28, 2026

Copy link
Copy Markdown
Member Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #75514 [ run ] triggered by Bot. Commit: c26cbfd Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #75514 [ run ] completed with state SUCCESS. Commit: c26cbfd
/LLM/main/L0_MergeRequest_PR pipeline #62245 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@lowsfer
lowsfer force-pushed the kvcm2-remove-python-backend branch from c26cbfd to 83cb8a1 Compare September 28, 2026 16:36
@lowsfer

lowsfer commented Sep 28, 2026

Copy link
Copy Markdown
Member Author

/bot run --disable-fail-fast

1 similar comment
@lowsfer

lowsfer commented Sep 28, 2026

Copy link
Copy Markdown
Member Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #75528 [ run ] triggered by Bot. Commit: 83cb8a1 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #75529 [ run ] triggered by Bot. Commit: 83cb8a1 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github/19154-83cb8a1 #75528 was force-killed by a newer pipeline run.
L0 job information not available (job may not have been triggered yet).

Link to superseding invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #75529 [ run ] completed with state SUCCESS. Commit: 83cb8a1
/LLM/main/L0_MergeRequest_PR pipeline #62258 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@lowsfer

lowsfer commented Sep 29, 2026

Copy link
Copy Markdown
Member Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #75610 [ run ] triggered by Bot. Commit: 83cb8a1 Link to invocation

KVCacheManagerV2 shipped two implementations of the same subsystem behind
TLLM_KV_CACHE_MANAGER_V2_BACKEND. The C++ port is the default, is what CI
exercises, and is the only backend newer features support, so the Python
implementation was carrying duplicate block-key hashing, eviction and stats
logic that had to stay bit-identical to C++, plus a mypyc and rawref build
pipeline that existed only to make it fast enough to matter.

tensorrt_llm/runtime/kv_cache_manager_v2/ is now a re-export shim over the
nanobind module plus the _introspection dispatcher: 36 tracked files down to 4.

Consumers that reached into private submodules move to the package surface. The
KV-aware router's V2 hashing is expressed with the existing
sequence_to_blockchain_keys, since every caller chains from a reuse-scope root,
so v2_sha256_block_hasher is gone and no hashing binding was needed. Only the
native disaggregated bounce buffer needed something new: PooledPhysMemAllocator
and VirtMem wrap the existing cudaVirtMem, and are registered on the
_introspection submodule rather than the package surface because they carry no
stability promise.

Streaming KV events (kv_cache_config.kv_events_config) are dropped. The sink is
duck-typed Python and the C++ radix tree calls its sink natively, so there is no
live path; validate_streaming_support now rejects the config and points at the
buffered path via event_buffer_max_size. The interface is kept as a stub and its
tests are skipped rather than deleted.

The --mypyc flag, TRTLLM_ENABLE_MYPYC, setup_mypyc.py, the rawref C extension
and the setup.py packaging surgery they required are all removed.

Signed-off-by: Yao Yao <lowsfer@users.noreply.github.com>
…atistics

Iteration statistics are keyed by life cycle and report recurrent (SSM) page
movement alongside attention movement, while the global cache-hit counters stay
attention-only. That split had no test on the C++ side: the only coverage lived
in the Python backend's unit tests, which went away with the backend itself, and
no C++ test builds an SSM life cycle at all.

Add a hybrid attention + SSM fixture and three cases over it:

  - offload and onboard are reported for both life cycles, with byte counts
    matching each life cycle's slot size
  - an SSM onboard leaves allocTotalBlocks / allocNewBlocks to attention
  - host drops are reported for both life cycles

Each case was confirmed to fail when the life-cycle filters are restored in
KvCache::_recordMigratedSlots and _recordDroppedPages.

Signed-off-by: Yao Yao <lowsfer@users.noreply.github.com>
@lowsfer
lowsfer force-pushed the kvcm2-remove-python-backend branch from 83cb8a1 to 4acacda Compare September 29, 2026 05:49
@lowsfer

lowsfer commented Sep 29, 2026

Copy link
Copy Markdown
Member Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #75626 [ run ] triggered by Bot. Commit: 4acacda Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #75610 [ run ] completed with state ABORTED. Commit: 83cb8a1

Link to invocation

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants