-
Notifications
You must be signed in to change notification settings - Fork 22k
Pull requests: ggml-org/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
llama : fix K/V and recurrent state cleanup after failed restores
testing
Everything test related
#27530
opened Aug 22, 2026 by
CHIPMUNK-T0T
Contributor
Loading…
Thread swizzling in kernel_mul_mm (Metal) for better cache locality.
Apple Metal
https://en.wikipedia.org/wiki/Metal_(API)
ggml
changes relating to the ggml tensor library for machine learning
testing
Everything test related
vulkan: combine duplicated fastdiv functions, rename the one optimizing small divs
ggml
changes relating to the ggml tensor library for machine learning
Vulkan
Issues specific to the Vulkan backend
#27526
opened Aug 22, 2026 by
jeffbolznv
Contributor
Loading…
cuda : fuse RWKV7 recurrent input and state paths
CUDA
Related to the CUDA backend
ggml
changes relating to the ggml tensor library for machine learning
testing
Everything test related
#27523
opened Aug 21, 2026 by
123123213weqw
Contributor
•
Draft
common: handle empty forced_tokens gracefully in reasoning-budget sampler
testing
Everything test related
#27514
opened Aug 21, 2026 by
arnavahire19
•
Draft
mamba2 : Flatten in/out projections to dispatch GEMM instead of GEMV
model
Model specific
#27513
opened Aug 21, 2026 by
pskrunner14
Loading…
Quant: OCP FP8 E4M3 support
conversion
ggml
changes relating to the ggml tensor library for machine learning
testing
Everything test related
common: add json.h abstraction
documentation
Improvements or additions to documentation
jinja parser
Issues related to the jinja parser
server
testing
Everything test related
#27511
opened Aug 21, 2026 by
ngxson
Collaborator
Loading…
5 tasks done
Add Q2_K reordered MMVQ and ESIMD kernels
ggml
changes relating to the ggml tensor library for machine learning
SYCL
https://en.wikipedia.org/wiki/SYCL - GPU programming language
#27509
opened Aug 21, 2026 by
malsbat
Contributor
Loading…
hexagon: the ggml_hexagon_supported_mul_mat function has incorrect logic when checking quantization types
ggml
changes relating to the ggml tensor library for machine learning
Hexagon
#27502
opened Aug 21, 2026 by
zhouwg-jeffzhou
Loading…
mtmd: fix hanging with specific video-vision mtmd in Windows
documentation
Improvements or additions to documentation
mtmd
Related to multimodal functionality (video/image/audio)
#27500
opened Aug 21, 2026 by
craftingmod
Loading…
fit: also take into account n_streams
server
#27496
opened Aug 21, 2026 by
ngxson
Collaborator
Loading…
SVE 128 bit Implementation of gemm_q4_k_8x8_q8_k kernel
ggml
changes relating to the ggml tensor library for machine learning
#27491
opened Aug 21, 2026 by
anubhavfujitsu
Loading…
misc : prevent RAM peaking at model loading stage
#27483
opened Aug 21, 2026 by
tdakhran
Contributor
Loading…
ggml : speed up batch-1 CPU decode, align large allocations
ggml
changes relating to the ggml tensor library for machine learning
#27478
opened Aug 21, 2026 by
matevz-kovacic
Loading…
common: set Muse Glimmer thinking tags in common_chat_params
#27475
opened Aug 21, 2026 by
paralin
Loading…
SVE 256 bit implementation of gemm_q6_K_8x8_q8_K kernel
ggml
changes relating to the ggml tensor library for machine learning
#27472
opened Aug 21, 2026 by
anubhavfujitsu
Loading…
Previous Next
ProTip!
Find all pull requests that aren't related to any open issues with -linked:issue.