Repository navigation
cuda: take the planar PTQ1_0 layout at one column on Turing as well as Ampere - #325
Open
professorpalmer wants to merge 1 commit into
Open
professorpalmer wants to merge 1 commit into
professorpalmer wants to merge 1 commit into
Conversation
…s Ampere One-column PTQ1_0 decode kept the Ada SOA_ISUM layout on sm_75. The planar PT kernel from PrismML-Eng#218 is faster there, as it is on Ampere: RTX 2060 SUPER, Bonsai 2 27B, q8_0 K/V, 4k context, greedy, 2 prompts x 400 tokens x 2 reps, stock clocks: SOA_ISUM 38.07 tok/s, PT 43.34 / 43.22 (+13.7%). GGML_CUDA_BATCH_INVARIANT=1 output is byte-identical before and after the change.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
One-line change: one-column PTQ1_0 decode on Turing (sm_75) takes the planar PT layout from #218, as Ampere already does. Single commit on current
prism(6bfcd79), applies without conflicts. Ada and newer keep SOA_ISUM; Ampere is unchanged.What it does
ggml_cuda_q8_1_layout_hostsends the one-column case toGGML_CUDA_Q8_1_PTon NVIDIA cc 800-889 and leaves every other NVIDIA card on the Ada-tunedSOA_ISUM. Turing (cc 750) fell through to SOA_ISUM. The condition now starts atGGML_CUDA_CC_TURINGinstead ofGGML_CUDA_CC_AMPERE. No kernel changes.Receipts
RTX 2060 SUPER 8 GB (sm_75, PCIe 3.0, stock clocks, power limit 175 W verified before every load), Windows, CUDA 13.4, Bonsai 2 27B
Ternary-Bonsai-2-27B-PTQ1_0.gguf, q8_0 K/V, 4k context, greedy, 2 prompts x 400 tokens x 2 reps per arm, arms alternated, server restarted per arm:GGML_CUDA_BATCH_INVARIANT=1GGML_CUDA_BATCH_INVARIANT=1+13.7% decode. An earlier alternated pair on the same card (SOA 36.76 / 36.69 vs PT via
BATCH_INVARIANT41.23 / 41.25) is what pointed at the gate.test-backend-ops -o MUL_MAT -b CUDA0on sm_75: 1516/1516, default and withGGML_CUDA_BATCH_INVARIANT=1.GGML_CUDA_BATCH_INVARIANT=1the greedy output is byte-identical before and after the change (both already took PT).llama-perplexity -b 1 -ub 1 -c 2048 --chunks 4, f16 K/V, base = SOA_ISUM): mean KL 0.000146, top-1 agreement 99.39 +/- 0.12%, max KL 0.047. The PT path thatGGML_CUDA_BATCH_INVARIANT=1already selects measures 0.000158 / 99.49% against the same base, so this is the usual near-tie disagreement between two summation orders. For scale, q8_0 vs f16 K/V on this card is mean KL 0.000167, top-1 99.42%.Thanks to @sudoingX for the PT kernel (#218); this only widens where it is used.