Skip to content

PTQ1_0 generation is ~9x slower on Vulkan than CUDA on the same GPU; Q4_K_M on the same card is 0.80x #201

Description

@blanham

Summary

On one RTX 2070 Max-Q (Turing, sm_75), PTQ1_0 token generation runs at 0.11× the CUDA rate under Vulkan, while an ordinary Q4_K_M model on the same card in the same session runs at 0.80×. Prompt processing is only 1.67× apart. That pattern points at the ternary mat-vec path in ggml-vulkan specifically, rather than at the Vulkan backend, the card, or the PTQ1_0 format.

PTQ1_0 on CUDA looks healthy — it scales from the Q4_K_M control by parameter count with nothing unexplained.

Version

prism-b10685-7dffb15 (7dffb158d)

Environment

  • GPU: NVIDIA GeForce RTX 2070 with Max-Q Design (TU106M), compute capability 7.5
  • Driver: 580.178.04, nvidia-open-driver-G06
  • OS: openSUSE Tumbleweed
  • Vulkan build: on host, GGML_VULKAN=ON
  • CUDA build: in a container, GGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=75, CUDA 12.8.1, linked against the driver stub and run against the host driver
  • Models: Ternary-Bonsai-2-27B-PTQ1_0.gguf (5,946,648,928 bytes), Qwen2.5-7B-Instruct-Q4_K_M.gguf

The machine also has an Intel UHD iGPU, so the discrete card is Vulkan1, not Vulkan0. All Vulkan rows below pass --device Vulkan1 explicitly.

Measurements

llama-bench -ngl 99 -p 128 -n 32 -r 2

model backend device pp128 t/s tg32 t/s
Bonsai 2 27B PTQ1_0 CUDA 2070 185.62 ± 3.19 16.09 ± 0.01
Bonsai 2 27B PTQ1_0 Vulkan 2070 110.90 ± 0.13 1.81 ± 0.02
Qwen2.5-7B Q4_K_M CUDA 2070 1252.45 ± 3.65 55.53 ± 0.16
Qwen2.5-7B Q4_K_M Vulkan 2070 1053.31 ± 195.26 44.60 ± 0.14

Ratios, Vulkan ÷ CUDA on the same card:

prompt (pp128) generation (tg32)
Q4_K_M 0.84× 0.80×
PTQ1_0 0.60× 0.11×

Why this looks like the mat-vec path

  1. It is not the backend. Vulkan reaches 80% of CUDA on Q4_K_M on this card.
  2. It is not the card. Same GPU, same driver, same session for every row.
  3. It is not the format. PTQ1_0 on CUDA is 16.09 t/s, and the control predicts it: 26.90 B params ÷ 7.62 B = 3.53× the compute per token, and 55.53 ÷ 3.53 = 15.7 against 16.09 measured.
  4. Prompt vs generation splits the way a mat-vec regression would. Batched matmul is 1.67× apart; mat-vec is 8.9× apart.

The device reports matrix cores: NV_coopmat2 and int dot: 1, so coopmat2 is available here.

Repro

# discrete card is Vulkan1 on this box; check with --list-devices
./llama-bench -m Ternary-Bonsai-2-27B-PTQ1_0.gguf --device Vulkan1 -ngl 99 -p 128 -n 32 -r 2
./llama-bench -m Qwen2.5-7B-Instruct-Q4_K_M.gguf   --device Vulkan1 -ngl 99 -p 128 -n 32 -r 2

Compare each against the same command on a CUDA build of the same commit.

Caveats

This is an observation, not a diagnosis — I have not read the shader. -r 2, one machine, one session, short context. The Q4_K_M control is one model at one quant, so "Vulkan is healthy on ordinary quants" is scoped to that row rather than claimed generally.

One incidental note that may help others comparing backends: llama-bench loads its ggml backend from the library path, so a Vulkan binary with a CUDA build's directory on LD_LIBRARY_PATH will silently run CUDA and report backend CUDA. Worth checking the backend column when two builds coexist.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions