Skip to content

Kossakowski-form dissipator - #222

Merged
david-pl merged 8 commits into
split/3-symmetric-evolutionfrom
split/5-kossakowski
Sep 28, 2026
Merged

david-pl merged 8 commits into
split/3-symmetric-evolutionfrom
split/5-kossakowski

Conversation

@david-pl

Copy link
Copy Markdown
Collaborator

Adds the general GKSL form
D*(O) = Σ_nm K_nm (A_n† O A_m − ½{A_n† A_m, O}) to ppvm-lindblad,
alongside the existing jump-operator form, with Python bindings, tests
and benchmarks.

Replaces #179, which was stacked on the pre-split #178 and could not be
reopened usefully: its branch had never been rebased, so its diff against
the split stack was 160 files and reverted work that #180–#182 had since
landed. This branch is the same feature ported onto split/3 and cleaned
up; the original commits are preserved on kossakowski-dissipator.

Stacked on #182 — review after that one settles, see the note below.

What's here

Commit
7b52cf6b core dissipator: kossakowski.rs, PairShape, add_kossakowski
e6266b52 superradiance + dipolar-relaxation benchmarks, profiling example
dc8c6234 Lindbladian(..., kossakowski=(ops, K)) Python surface
b52d22cc dense-reference and orbit-representative tests

Conjugate pairs (n,m)/(m,n) are folded into a single upper-triangle
entry — both sandwiches produce the same output words with conjugate
phase and the action keeps only the real part — halving the pair count.
PairShape distinguishes the diagonal entry from a folded off-diagonal
one and carries the term list only the folded case needs, so the
off-diagonal-only path is unreachable on a diagonal pair.

Against the equivalent eigenmode-jump decomposition on a superradiant
chain, the pair-matrix path costs 27.7 ms vs 103 ms at n=10 and 96 ms vs
1.71 s at n=30: the per-string action scales with the number of nonzero
K_nm entries instead of carrying an extra factor of M.

Differences from #179

Ported into the spec/basis/step/algebra layout split/2 and
split/3 introduced, rather than appended to lib.rs:

  • New kossakowski.rs module. spec.rs grows 65 lines, not 520.
  • add_kossakowski (147 lines) split into validate_k / parse_ops /
    compile_pair / compile; the ~90-line action arm became
    Pair::accumulate + the one-sided and both-sided halves.
  • PairShape enum replaces an off_diag bool that drove three separate
    branches in the hot path. Benchmarked both shapes: the deltas
    (+2.4% / −1.3% / +2.2%, n30 p=0.17) are smaller than the drift on the
    eigenmode_* control benchmarks (+3.4% to +6.6%) in the same run, so
    this is noise, not a regression.
  • 7 clippy allows removed, 0 added. split/3 had none in
    ppvm-lindblad; three needless_range_loop became iterator loops and
    four type_complexity became two named aliases.
  • precompute_ldagger_l generalized to precompute_adag_b(a, b), moved
    to algebra.rs beside PauliTerm and a named COEFF_DROP_TOL (was a
    1e-14 literal in five places).
  • Dropped a dead let _ = i;, and row-length errors now report which row.
  • support_mask lifted out of the pair loop, where it was being rebuilt
    per iteration.
  • Deleted the k_ops field — LindbladSpec state that never outlived
    the constructor.
  • Did not port the original's :func:/:meth: docstrings; split/3
    had already converted those to backticks per AGENTS.md.

One thing to check

test_orbit_dissipative.py failed on arrival. d8313e0e on split/3
("orbit-rep evolution on stabilized orbits") made rep coefficients member
coefficients on every orbit, where stabilized orbits were previously
divided by |G|; the tests encoded the old behaviour and c_I came out
exactly 6× the asserted value on a 6-site ring.

I rewrote the reconstruction as Σ |orbit(rep)| · c_rep and updated the
identity-bookkeeping expectation to coeff_I. The new values match the
closed form exactly. If d8313e0e was not meant to change that
convention, the bug is there and not in these tests
— worth confirming
while reviewing #182.

Verification

  • cargo test -p ppvm-lindblad — 8 passed
  • pytest test/lindblad/ — 35 passed, 1 skipped
  • cargo clippy -p ppvm-lindblad --all-targets — zero warnings
  • pre-commit green on all four commits

🤖 Generated with Claude Code

david-pl and others added 4 commits September 16, 2026 09:58
Add the general GKSL form
`D*(O) = Σ_nm K_nm (A_n† O A_m − ½{A_n† A_m, O})`, alongside the existing
jump-operator form. Each `(n, m)` entry is compiled once into a
`kossakowski::Pair`; Hermitian-conjugate pairs are folded into a single
upper-triangle entry, since both sandwiches produce the same output words
with conjugate phase and the action keeps only the real part.

`PairShape` distinguishes the diagonal `(n,n)` entry from a folded
off-diagonal one and carries the term list only the folded case needs, so
the off-diagonal-only path cannot be reached on a diagonal pair.

`precompute_ldagger_l` generalizes to `precompute_adag_b(a, b)` and moves
to `algebra` next to `PauliTerm`, which the new module shares.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…kloads

`kossakowski` compares the pair-matrix path against the equivalent
eigenmode-jump decomposition on a superradiant chain; `drug_dipolar`
covers the 2-local rank-2 tensors of dipolar relaxation, where each `A_n`
carries several Pauli terms. `drug_profile` is the same model wired to a
plain timing loop for use under a profiler.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`Lindbladian(..., kossakowski=(ops, K))` forwards an operator family and
its pair matrix to the native spec. Mathematically identical to passing
the eigenmode jumps of `K = V diag(γ) V†`, but the per-string action cost
scales with the number of nonzero `K_nm` entries rather than carrying an
extra factor of `M`.

Python validates `K` for Hermiticity and positive-semidefiniteness before
the call. The core rejects a non-Hermitian `K` too, but the Python check
is not redundant: `eigvalsh` reads only one triangle, so without it a
non-Hermitian `K` would pass the PSD test on its symmetric part.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`test_kossakowski` checks the pair-matrix action against a dense
superoperator and against the equivalent eigenmode-jump decomposition;
`test_orbit_dissipative` covers the orbit-representative path on a ring
of emitters, including the non-unital Z → I flow across orbits of
different sizes.

The orbit tests reconstruct a translation-invariant sum as
`Σ |orbit(rep)| · c_rep`. Rep coefficients are member coefficients on
every orbit since the stabilized-orbit fix in
d8313e0, so the previous uniform `|G| ·` factor would over-count orbits
with a non-trivial stabilizer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-09-28 13:46 UTC

david-pl and others added 4 commits September 28, 2026 13:53
`LindbladSpec` was hard-capped at 128 qubits by the compile-time
`W_CHUNKS`, so e.g. a 130-qubit Lindbladian failed with

    ValueError: LindbladSpec supports n_qubits ≤ 128; got 130

which `wide-pauli-words` (4c07999) had already fixed; the fix did not
make it into the split. Port it onto the current spec/basis/step layout,
including the Kossakowski dissipator added in this PR.

- `ppvm-lindblad` is const-generic in the chunk count: `Word<C>`,
  `LindbladSpec<C>`, and every function touching words. Both default to
  `C = W_CHUNKS` (128 qubits), so existing Rust callers are unchanged.
  `max_qubits::<C>()`, `chunks_for(n)`, `WIDTHS` (128/256/512-qubit chunk
  counts, in `Chunk` units so the wasm32 u32 layout keeps working) and
  `MAX_SUPPORTED_QUBITS = 512` are exported. `Error::TooManyQubits` now
  carries the width's capacity.
- PyO3: `LindbladSpec` holds an `AnySpec { W128, W256, W512 }` picked at
  construction and dispatches with `with_spec!`; the symmetry bindings use
  `with_width!`. Registers of ≤128 qubits monomorphize to exactly the old
  layout. More than 512 qubits is a `ValueError`.
- The orbit path's masked-shift generators cover up to 1024 bits, so all
  three widths take the fast canonicalizer.

Tests: `tests/word_width.rs` — capacities, the 130/512-qubit cases, the
width error, bit-identical real-path results for the same problem padded
into 128/256/512-qubit words (Kossakowski and orbit-rep to rounding: their
accumulation order follows word hashes, which depend on the width), and
an orbit-rep step at 130 qubits. Python: `test_word_width.py` builds and
steps 130–512-qubit Lindbladians, rejects 513, and runs the momentum-orbit
step across the 128-qubit boundary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TmnGNkxQqpmFmGszAV7dm5
)

Stacked on #222 (`split/5-kossakowski`).

## Summary

`build_mf_cols` / `build_orbit_rep_cols` returned one `Vec` per column,
each reserved at the full L* output count although only in-basis entries
are kept. The expm generator cache is now a `BlockCsc`: blocks of up to
4096 columns, each an exactly sized `offsets`/`rows`/`vals` triple
(structure of arrays, 12 B per real entry instead of 16 B). `CscOp::dot`
keeps the same per-thread column partition, so results are
bit-identical.

Single commit, `crates/ppvm-lindblad/src/mf_expm.rs` only.

## Measurements

`pc_step` on the two-leg XY ladder with a local probe (not part of this
PR): L = 41 rungs (N = 82), `max_basis = 2^18`, `admit_basis = 3·2^18`,
dt = 0.1, 25 steps, 4 threads, `drop_tol = 0`. Apple M4 MacBook Air
(fanless, so wall times are only comparable within a rep). Max RSS from
`/usr/bin/time -l`; live-heap peak from a counting global allocator.

| mimalloc, 3 interleaved reps | wall (s) | max RSS (MiB) | live-heap
peak (MiB) |
|---|---|---|---|
| base (#222) | 53.7 / 53.5 / 62.5 | 484 / 484 / 483 | 441.7 |
| this PR | 46.2 / 48.9 / 60.1 | 501 / 491 / 470 | 275.2 |

- The live-heap peak drops by 38% and wall time by 4–14%, with each expm
call about 20–30% faster.
- On macOS + mimalloc the heap saving shows up only weakly in the
process footprint: max RSS moves by −3% to +4%, and peak footprint
(`/usr/bin/time -l`) drops by about 3–13%. With the macOS system
allocator, RSS is not reproducible (768–1237 MiB for the same binary),
so the effect cannot be resolved there.
- This has not been measured on Linux/glibc, where fragmentation from
the per-column allocations of the old layout may matter more.
- The final coefficients are bit-identical to the base in every run.

Tests: `cargo test -p ppvm-lindblad` passes (8/8) with this change
stacked together with the skip-empty-hop2 PR; not yet run on this branch
alone (CI will).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: David Plankensteiner <dplankensteiner@quera.com>
…othing (#230)

Stacked on #222 (`split/5-kossakowski`).

## Summary

With `admit_basis` set, once the basis has reached `max_basis` the first
leakage admission fills the whole room `admit − |basis|`. The second
leakage call then ran with room = 0 and admitted nothing in every capped
step, while still evaluating all its candidates, and the corrector
repeated the predictor exactly on the same basis and input.

`pc_step` and `pc_step_orbit_rep` now skip the second leakage pass when
room = 0 and reuse the predicted state whenever the second hop admitted
no string. Results are bit-identical; skipped phases report 0 in
`PcStepTimings`. Default behaviour is otherwise unchanged: capped steps
still carry no second-order admission, as before.

Single commit, `crates/ppvm-lindblad/src/step.rs` only.

## Measurements

`pc_step` on the two-leg XY ladder with a local probe (not part of this
PR): L = 41 rungs (N = 82), `max_basis = 2^18`, `admit_basis = 3·2^18`,
dt = 0.1, 25 steps, 4 threads, `drop_tol = 0`. Apple M4 MacBook Air
(fanless, so wall times are only comparable within a rep).

In the base, 21 of the 25 steps are capped and the second hop (leakage2
+ expm2) accounts for 53% of the wall time (34.5 of 65.2 s).

| | wall (s), per rep | 
|---|---|
| base (#222), mimalloc | 53.7 / 53.5 / 62.5 |
| this PR, mimalloc | 26.6 / 28.7 / 33.3 |
| base (#222), system malloc | 65.6 / 79.1 |
| this PR, system malloc | 32.4 / 38.4 |

- The run is 1.9–2.0× faster end-to-end; memory is unchanged (live-heap
peak 441.7 MiB in both).
- The final coefficients are bit-identical to the base in every run.

Tests: `cargo test -p ppvm-lindblad` passes (8/8) with this change
stacked together with the blocked-CSC PR; not yet run on this branch
alone (CI will).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: David Plankensteiner <david-pl@users.noreply.github.com>
@david-pl
david-pl merged commit 5f58273 into split/3-symmetric-evolution Sep 28, 2026
3 checks passed
@david-pl
david-pl deleted the split/5-kossakowski branch September 28, 2026 13:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants