Summary
On one RTX 2070 Max-Q (Turing, sm_75), PTQ1_0 token generation runs at 0.11× the CUDA rate under Vulkan, while an ordinary Q4_K_M model on the same card in the same session runs at 0.80×. Prompt processing is only 1.67× apart. That pattern points at the ternary mat-vec path in ggml-vulkan specifically, rather than at the Vulkan backend, the card, or the PTQ1_0 format.
PTQ1_0 on CUDA looks healthy — it scales from the Q4_K_M control by parameter count with nothing unexplained.
Version
prism-b10685-7dffb15 (7dffb158d)
Environment
- GPU: NVIDIA GeForce RTX 2070 with Max-Q Design (TU106M), compute capability 7.5
- Driver: 580.178.04,
nvidia-open-driver-G06
- OS: openSUSE Tumbleweed
- Vulkan build: on host,
GGML_VULKAN=ON
- CUDA build: in a container,
GGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=75, CUDA 12.8.1, linked against the driver stub and run against the host driver
- Models:
Ternary-Bonsai-2-27B-PTQ1_0.gguf (5,946,648,928 bytes), Qwen2.5-7B-Instruct-Q4_K_M.gguf
The machine also has an Intel UHD iGPU, so the discrete card is Vulkan1, not Vulkan0. All Vulkan rows below pass --device Vulkan1 explicitly.
Measurements
llama-bench -ngl 99 -p 128 -n 32 -r 2
| model |
backend |
device |
pp128 t/s |
tg32 t/s |
Bonsai 2 27B PTQ1_0 |
CUDA |
2070 |
185.62 ± 3.19 |
16.09 ± 0.01 |
Bonsai 2 27B PTQ1_0 |
Vulkan |
2070 |
110.90 ± 0.13 |
1.81 ± 0.02 |
Qwen2.5-7B Q4_K_M |
CUDA |
2070 |
1252.45 ± 3.65 |
55.53 ± 0.16 |
Qwen2.5-7B Q4_K_M |
Vulkan |
2070 |
1053.31 ± 195.26 |
44.60 ± 0.14 |
Ratios, Vulkan ÷ CUDA on the same card:
|
prompt (pp128) |
generation (tg32) |
Q4_K_M |
0.84× |
0.80× |
PTQ1_0 |
0.60× |
0.11× |
Why this looks like the mat-vec path
- It is not the backend. Vulkan reaches 80% of CUDA on
Q4_K_M on this card.
- It is not the card. Same GPU, same driver, same session for every row.
- It is not the format.
PTQ1_0 on CUDA is 16.09 t/s, and the control predicts it: 26.90 B params ÷ 7.62 B = 3.53× the compute per token, and 55.53 ÷ 3.53 = 15.7 against 16.09 measured.
- Prompt vs generation splits the way a mat-vec regression would. Batched matmul is 1.67× apart; mat-vec is 8.9× apart.
The device reports matrix cores: NV_coopmat2 and int dot: 1, so coopmat2 is available here.
Repro
# discrete card is Vulkan1 on this box; check with --list-devices
./llama-bench -m Ternary-Bonsai-2-27B-PTQ1_0.gguf --device Vulkan1 -ngl 99 -p 128 -n 32 -r 2
./llama-bench -m Qwen2.5-7B-Instruct-Q4_K_M.gguf --device Vulkan1 -ngl 99 -p 128 -n 32 -r 2
Compare each against the same command on a CUDA build of the same commit.
Caveats
This is an observation, not a diagnosis — I have not read the shader. -r 2, one machine, one session, short context. The Q4_K_M control is one model at one quant, so "Vulkan is healthy on ordinary quants" is scoped to that row rather than claimed generally.
One incidental note that may help others comparing backends: llama-bench loads its ggml backend from the library path, so a Vulkan binary with a CUDA build's directory on LD_LIBRARY_PATH will silently run CUDA and report backend CUDA. Worth checking the backend column when two builds coexist.
Summary
On one RTX 2070 Max-Q (Turing,
sm_75),PTQ1_0token generation runs at 0.11× the CUDA rate under Vulkan, while an ordinaryQ4_K_Mmodel on the same card in the same session runs at 0.80×. Prompt processing is only 1.67× apart. That pattern points at the ternary mat-vec path inggml-vulkanspecifically, rather than at the Vulkan backend, the card, or thePTQ1_0format.PTQ1_0on CUDA looks healthy — it scales from theQ4_K_Mcontrol by parameter count with nothing unexplained.Version
prism-b10685-7dffb15(7dffb158d)Environment
nvidia-open-driver-G06GGML_VULKAN=ONGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=75, CUDA 12.8.1, linked against the driver stub and run against the host driverTernary-Bonsai-2-27B-PTQ1_0.gguf(5,946,648,928 bytes),Qwen2.5-7B-Instruct-Q4_K_M.ggufThe machine also has an Intel UHD iGPU, so the discrete card is
Vulkan1, notVulkan0. All Vulkan rows below pass--device Vulkan1explicitly.Measurements
llama-bench -ngl 99 -p 128 -n 32 -r 2PTQ1_0PTQ1_0Q4_K_MQ4_K_MRatios, Vulkan ÷ CUDA on the same card:
Q4_K_MPTQ1_0Why this looks like the mat-vec path
Q4_K_Mon this card.PTQ1_0on CUDA is 16.09 t/s, and the control predicts it: 26.90 B params ÷ 7.62 B = 3.53× the compute per token, and 55.53 ÷ 3.53 = 15.7 against 16.09 measured.The device reports
matrix cores: NV_coopmat2andint dot: 1, so coopmat2 is available here.Repro
# discrete card is Vulkan1 on this box; check with --list-devices ./llama-bench -m Ternary-Bonsai-2-27B-PTQ1_0.gguf --device Vulkan1 -ngl 99 -p 128 -n 32 -r 2 ./llama-bench -m Qwen2.5-7B-Instruct-Q4_K_M.gguf --device Vulkan1 -ngl 99 -p 128 -n 32 -r 2Compare each against the same command on a CUDA build of the same commit.
Caveats
This is an observation, not a diagnosis — I have not read the shader.
-r 2, one machine, one session, short context. TheQ4_K_Mcontrol is one model at one quant, so "Vulkan is healthy on ordinary quants" is scoped to that row rather than claimed generally.One incidental note that may help others comparing backends:
llama-benchloads its ggml backend from the library path, so a Vulkan binary with a CUDA build's directory onLD_LIBRARY_PATHwill silently run CUDA and reportbackend CUDA. Worth checking thebackendcolumn when two builds coexist.