This fork (branch ktrain) is PrismML's prism plus nine pull requests that speed up decoding of the ternary Bonsai 2 27B models (PTQ1_0 / PQ2_0, with an MTP draft head) on NVIDIA Ada (RTX 4070, cc 8.9). They are offered to PrismML as separate pull requests; this branch merges all of them for people who do not want to wait for the merges. Every patch is also a branch of its own off prism (pr/*).
Speed against prism (2459f68b5, RTX 4070 12 GB, Ternary-Bonsai-2-27B PTQ1_0 + MTP head, q4_0 K/V, greedy, same build settings):
- Agent setup (MTP
n-max 2, context 114688): 104.2 -> 110.6 tok/s, +5.9 % (95 % CI [+5.7, +6.1] %, 8 interleaved pairs, outputs identical). - 120k tokens of filled context, no MTP: 25.5 -> 46.9 tok/s, +84 % (two runs per build).
| PR | Branch | What | Effect on the RTX 4070 |
|---|---|---|---|
| #306 | pr/pq2_0-multicol |
PQ2_0 mat-vec kernel for 3-8 columns | +17 / +25 / +29 % decode with 4 / 6 / 8 parallel sequences |
| #307 | pr/fattn-gqa-mma |
MMA flash attention for GQA > 4 with quantized K/V | 120k context, no MTP: 25.3 -> 41.2 tok/s |
| #308 | pr/fattn-mma-tile |
smaller KV tile for the 8-column MMA config (head size 256) | 120k context, no MTP: 41.2 -> 46.0 tok/s (with #307) |
| #309 | pr/concat-transpose-all-gpus |
transposing concat kernel on all GPUs | +1.2 % decode with MTP |
| #310 | pr/ptq1-gate-up-fuse-mc |
gate + up + SwiGLU fused PTQ1_0 mat-vec for 2-4 columns | +1.3 % decode with MTP |
| #311 | pr/kv-seq-rm-bound |
seq_rm only over the used cell range |
+0.8 % at 114688 context |
| #312 | pr/server-ckpt-buffer-reuse |
reuse the buffers of evicted prompt checkpoints | -20 ms per turn in a long multi-turn chat |
| #313 | pr/sampler-topk-from-logits |
top-k straight from the logits | +1.9 % decode with MTP |
| #314 | pr/ptq1-l2-prefetch |
prefetch the next PTQ1_0 mat-vec's weights into L2 | +2.0 % decode with MTP |
#306, #307 and #308 change the arithmetic order, so results differ at rounding level; the other six give identical outputs. Everything was measured on one GPU and one model family only; other hardware is untested. Details, measurement method and limits: docs/ktrain/README.md.
MTP version of the PTQ1_0 model: docs/ktrain/mtp/README.md shows how to build Ternary-Bonsai-2-27B-PTQ1_0-MTP-Q8_0.gguf from the two published files.
AI assistance: the patches were developed with Claude Code; see docs/ktrain/README.md.
Important
This is the PrismML fork of llama.cpp, the main line behind the Bonsai models (branch prism, developed as prism-v7). It tracks current mainline llama.cpp and adds the fork's low-bit formats and runtime features on top.
New here? Start with the Bonsai-demo repo. It downloads the right models and the correct prebuilt binaries for your hardware/backend automatically.
Which ternary model file to use:
*-PQ2_0.gguf(fork group-128, ggml id 142): preferred on Metal, CUDA, HIP and CPU. About 6% smaller than group-64.*-Q2_0_g64.gguf/ 27B*-Q2_g64.gguf(official group-64, ggml id 42): runs on every backend here AND on mainline llama.cpp. If unsure, use this. Newer model releases name this file plain*-Q2_0.gguf.*-Q2_0.ggufon OLDER model repos is the deprecated legacy format (group 128 stored as id 42). It does not load on these builds; the error tells you which file to get instead. If you must run it, use the frozenprism-v5line and its final releaseprism-b9601.
Speculative decoding (dspark) is supported via mainline's draft-dspark plus fork patches. Drafters published for older model releases need a one-time conversion with gguf-dspark-to-dflash (see SPECULATIVE.md in Bonsai-demo); newer releases ship ready-to-use drafters.
Do NOT build from prism-v6 (stale mid-migration snapshot) and do NOT mix this fork's ggml-* libraries with a stock llama.cpp build.
LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
A few options to get llama.cpp installed on your machine:
- Visit https://llama.app and follow the instructions
- Run with Docker - see our Docker documentation
- Download pre-built binaries from the releases page
- Build from source by cloning this repository - check out our build guide
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
VLM session with llama cli
|
Built-in web UI against llama serve
|
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
- Plain C/C++ implementation without any dependencies
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
- AVX, AVX2, AVX512 and AMX support for x86 architectures
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon [In Progress] | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
- How to build
- Running on Docker
- Build on Android
- Multi-GPU usage
- Performance troubleshooting
- GGML tips & tricks
- XCFramework
- Completions
- Models
- Release process
- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the
llama.cpprepo and merge PRs into themasterbranch - Any help with managing issues, PRs and projects is very appreciated!
- Read the CONTRIBUTING.md for more information
- yhirose/cpp-httplib - Single-header HTTP server, used by
llama-server- MIT license - nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
- nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
- mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
- sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain

