Skip to content

Latest commit

 

History

10,199 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

llama.cpp — DualDeadline branch

Paper on ResearchGate

Note

Looking for TurboQuant? The 3-bit KV-cache quantization work this repository is named after lives on the main branch. This branch (llama-dualdeadline) is an independent research prototype.

This branch implements DualDeadline — component-staged expert streaming for offloaded mixture-of-experts inference, after the paper DualDeadline: Component-Staged Expert Prefetching for Exact Offloaded Mixture-of-Experts Inference (Marcel Butucea, 2026; study artifacts).

A gated MoE expert has two transfer deadlines, not one: gate/up projections are needed when expert compute starts, but the down projection is not consumed until gate/up arithmetic finishes. When expert tensors are kept in pinned host RAM (--cpu-moe --no-mmap), this branch lets the GPU claim the batch-1 expert matmuls by copying only the router-selected experts' slices into a persistent per-tensor LRU cache — gate/up slices first, down slices overlapping gate/up compute on a dedicated copy stream. Routing and outputs are exact; there is no prediction and no approximation.

GGML_CUDA_DUAL_DEADLINE=1 \
llama-cli -m model.gguf -ngl 999 --cpu-moe --no-mmap -p "..."

Measured on Qwen3.6-35B-A3B (Q4_K_XL, 256 experts/layer, top-8) with all expert tensors in host memory, AMD Ryzen AI MAX+ 395 / Radeon 8060S (Strix Halo, unified memory), ROCm 7.2.4, llama-bench -r 3, tg64:

configuration tg64 t/s
upstream: expert matmuls on CPU (16 threads) 49.7 ± 0.3
DualDeadline staged, cache 64 experts/tensor (84% hit rate) 44.1 ± 0.4
DualDeadline monolithic, cache 64 (same copies, single deadline) 42.7 ± 0.4
DualDeadline staged, cache 32 41.6 ± 0.4
DualDeadline monolithic, cache 32 39.6 ± 0.4

Two honest readings. First, staged beats monolithic by 3–5% end-to-end at identical bytes, cache, and kernels — that is the paper's two-deadline claim reproduced inside a real decode loop rather than a microbenchmark. Second, on this unified-memory APU the CPU baseline still wins: there is no PCIe link to hide transfers behind, so explicit staging competes with zero-cost DRAM reads. The regime this mechanism targets is a PCIe-attached discrete GPU whose VRAM cannot hold the experts — where the alternative to staging is CPU compute or full-tensor streaming — and that measurement is still to be done. A validation mode (GGML_CUDA_DD_VALIDATE=1) recomputes every intercepted matmul via the standard path: 3,120 ops checked, max abs err 0.0 — bitwise identical.

See docs/dual-deadline.md for the design, all environment variables, limitations, and how this relates to upstream's used-expert streaming. Implementation: ggml/src/ggml-cuda/dual-deadline.cu.


llama

Quick start

A few options to get llama.cpp installed on your machine:

Once installed:

# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF

# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
VLM session with `llama cli` VLM session with llama cli Built-in web UI against `llama serve` running Qwen 3.6 Built-in web UI against llama serve

Description

The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud.

  • Plain C/C++ implementation without any dependencies
  • Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
  • AVX, AVX2, AVX512 and AMX support for x86 architectures
  • RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
  • 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
  • Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
  • Vulkan and SYCL backend support
  • CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity

The llama.cpp project is build on top of the ggml library.

Supported backends

Backend Target devices
BLAS All
BLIS All
CANN Ascend NPU
CUDA Nvidia GPU
HIP AMD GPU
Hexagon [In Progress] Snapdragon
IBM zDNN IBM Z & LinuxONE
MUSA Moore Threads GPU
Metal Apple Silicon
OpenCL Adreno GPU
OpenVINO [In Progress] Intel CPUs, GPUs, and NPUs
RPC All
SYCL Intel GPU
VirtGPU VirtGPU APIR
Vulkan GPU
WebGPU All
ZenDNN AMD CPU

Documentation

Tools

Development

Contributing

  • Contributors can open PRs
  • Collaborators will be invited based on contributions
  • Maintainers can push to branches in the llama.cpp repo and merge PRs into the master branch
  • Any help with managing issues, PRs and projects is very appreciated!
  • Read the CONTRIBUTING.md for more information

Acknowledgements

  • yhirose/cpp-httplib - Single-header HTTP server, used by llama-server - MIT license
  • stb-image - Single-header image format decoder, used by multimodal subsystem - Public domain
  • nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
  • miniaudio.h - Single-header audio format decoder, used by multimodal subsystem - Public domain
  • subprocess.h - Single-header process launching solution for C and C++ - Public domain

About

TurboQuant 3-bit KV-cache quantization for llama.cpp

Topics

Resources

Contributing

Security policy

Stars

55 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages