Note
Looking for TurboQuant? The 3-bit KV-cache quantization work this
repository is named after lives on the main branch.
This branch (llama-dualdeadline) is an independent research prototype.
This branch implements DualDeadline — component-staged expert streaming for offloaded mixture-of-experts inference, after the paper DualDeadline: Component-Staged Expert Prefetching for Exact Offloaded Mixture-of-Experts Inference (Marcel Butucea, 2026; study artifacts).
A gated MoE expert has two transfer deadlines, not one: gate/up
projections are needed when expert compute starts, but the down projection is
not consumed until gate/up arithmetic finishes. When expert tensors are kept
in pinned host RAM (--cpu-moe --no-mmap), this branch lets the GPU claim the
batch-1 expert matmuls by copying only the router-selected experts' slices
into a persistent per-tensor LRU cache — gate/up slices first, down slices
overlapping gate/up compute on a dedicated copy stream. Routing and outputs
are exact; there is no prediction and no approximation.
GGML_CUDA_DUAL_DEADLINE=1 \
llama-cli -m model.gguf -ngl 999 --cpu-moe --no-mmap -p "..."Measured on Qwen3.6-35B-A3B (Q4_K_XL, 256 experts/layer, top-8) with all
expert tensors in host memory, AMD Ryzen AI MAX+ 395 / Radeon 8060S
(Strix Halo, unified memory), ROCm 7.2.4, llama-bench -r 3, tg64:
| configuration | tg64 t/s |
|---|---|
| upstream: expert matmuls on CPU (16 threads) | 49.7 ± 0.3 |
| DualDeadline staged, cache 64 experts/tensor (84% hit rate) | 44.1 ± 0.4 |
| DualDeadline monolithic, cache 64 (same copies, single deadline) | 42.7 ± 0.4 |
| DualDeadline staged, cache 32 | 41.6 ± 0.4 |
| DualDeadline monolithic, cache 32 | 39.6 ± 0.4 |
Two honest readings. First, staged beats monolithic by 3–5% end-to-end at
identical bytes, cache, and kernels — that is the paper's two-deadline claim
reproduced inside a real decode loop rather than a microbenchmark. Second, on
this unified-memory APU the CPU baseline still wins: there is no PCIe link
to hide transfers behind, so explicit staging competes with zero-cost DRAM
reads. The regime this mechanism targets is a PCIe-attached discrete GPU
whose VRAM cannot hold the experts — where the alternative to staging is CPU
compute or full-tensor streaming — and that measurement is still to be done.
A validation mode (GGML_CUDA_DD_VALIDATE=1) recomputes every intercepted
matmul via the standard path: 3,120 ops checked, max abs err 0.0 —
bitwise identical.
See docs/dual-deadline.md for the design, all
environment variables, limitations, and how this relates to upstream's
used-expert streaming. Implementation:
ggml/src/ggml-cuda/dual-deadline.cu.
LLM inference in C/C++
manifesto / ggml / ops / maintainer PRs / dev branches / compile times / lib llama API / llama-server REST API
A few options to get llama.cpp installed on your machine:
- Visit https://llama.app and follow the instructions
- Run with Docker - see our Docker documentation
- Download pre-built binaries from the releases page
- Build from source by cloning this repository - check out our build guide
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
VLM session with llama cli
|
Built-in web UI against llama serve
|
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
- Plain C/C++ implementation without any dependencies
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
- AVX, AVX2, AVX512 and AMX support for x86 architectures
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon [In Progress] | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
- How to build
- Running on Docker
- Build on Android
- Multi-GPU usage
- Performance troubleshooting
- GGML tips & tricks
- XCFramework
- Completions
- Models
- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the
llama.cpprepo and merge PRs into themasterbranch - Any help with managing issues, PRs and projects is very appreciated!
- Read the CONTRIBUTING.md for more information
- yhirose/cpp-httplib - Single-header HTTP server, used by
llama-server- MIT license - stb-image - Single-header image format decoder, used by multimodal subsystem - Public domain
- nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
- miniaudio.h - Single-header audio format decoder, used by multimodal subsystem - Public domain
- subprocess.h - Single-header process launching solution for C and C++ - Public domain

