forked from ggml-org/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 73
Pull requests: PrismML-Eng/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
ggml-cpu: add opt-in Q2_0 VNNI64 four-row decode
documentation
Improvements or additions to documentation
ggml
#95
opened Jul 20, 2026 by
chris-lee-mc
•
Draft
server: do not propagate --kv-mean-center bias to draft model contexts
#81
opened Jul 16, 2026 by
cnndabbler
Loading…
gb10-blackwell: env-gated Blackwell int8 MMA for Q1_0/Q2_0 weights on DGX Spark
CUDA
documentation
Improvements or additions to documentation
ggml
#79
opened Jul 16, 2026 by
sumergoconicio
Loading…
ggml-cpu: enable Q2_0 VNNI kernel on AVX-VNNI-only CPUs
ggml
#76
opened Jul 15, 2026 by
gondoi
Loading…
Q2_0 group 64: CUDA backend
Apple Metal
conversion
CUDA
documentation
Improvements or additions to documentation
ggml
Hexagon
jinja parser
model
OpenCL
server/ui
server
SYCL
testing
Vulkan
#43
opened Jun 10, 2026 by
khosravipasha
Collaborator
•
Draft
opencl: Q1_0 support first attempt
ggml
OpenCL
#25
opened Apr 15, 2026 by
khosravipasha
Collaborator
•
Draft
ProTip!
Type g p on any issue or pull request to go back to the pull request listing page.