Skip to content

ggml 0.26.0 - #315641

Merged
BrewTestBot merged 2 commits into
mainfrom
bump-ggml-0.26.0
Oct 5, 2026
Merged

BrewTestBot merged 2 commits into
mainfrom
bump-ggml-0.26.0

Conversation

@BrewTestBot

Copy link
Copy Markdown
Contributor

Created by brew bump


Created with brew bump-formula-pr.

  • TODO and FIXME comments have been checked.
release notes
## Overview

ggml v0.26.0 brings a new alloc_buffer_n/get_alloc_size_n buffer allocation API, sparse flash attention kernels on SYCL, Vulkan and Metal, and major lightning indexer improvements (halved score memory, tiling, MUSA support). The CPU backend gains BF16 ops and a tiled k-quant mul_mat, CUDA gains a model-driven W4A4 (NVFP4/MXFP4) mul_mat path plus MMVQ shared-expert fusion, and the Hexagon backend adds a sampler and more quant types. Model loading is faster with stricter GGUF size validation, Windows ARM64 MSVC builds are enabled, and WebGPU/OpenVINO/OpenCL/SYCL pick up numerous new kernels, ops and fixes.

API changes

  • Add alloc_buffer_n and get_alloc_size_n to the buffer type interface with public APIs ggml_backend_buft_alloc_buffer_n/ggml_backend_buft_get_alloc_size_n for multi-buffer allocation and size planning (llama/23671)
  • Add new public header include/ggml-zdnn.h for the zDNN backend (llama/29541)

Core changes

  • Speed up model loading and fix a hang on crafted GGUFs with very large KV dimensions (llama/29598)
  • Harden tensor/GGUF size validation: fix integer overflows and reject tensor sizes that wrap after padding (llama/29384, llama/26979)
  • Collect all input tensors into graph_inputs up front and require them to be GGML_OP_NONE, fixing spurious graph re-reserves with pipeline parallelism (llama/29634, llama/29647)
  • Meta backend: clear inactive AllReduce shards with FILL instead of SCALE (llama/29793)
  • ggml-quants: avoid invalid rounding in the qkx3 scale search (llama/29817)
  • Enable Windows ARM64 builds with MSVC cl.exe (llama/28362)

Backend changes

CPU

  • Support BF16/FP16/FP32 K tails in tinyBLAS on x86 (llama/29806)
  • BF16 op support: add unary, GLU, binary and scale ops (with CUDA) and accept BF16 in src1 of mul_mat (llama/29675, llama/28937)
  • Tiled mul_mat for k-quants via int8 unpack tiles and 16x16 microkernels (llama/27851)
  • Enable tiled flash attention for non-vector-multiple head dims on x86 (llama/29423)
  • Accumulate f16 dot products in f32 on AVX512-FP16 (llama/29545)
  • Add Q8_0 IME1 matrix kernel for SpacemiT X60 (llama/28479)
  • Fix soft_max_back wrong output when dst aliases src1 (llama/27096)

CUDA / HIP

  • Add model-driven W4A4 (NVFP4/MXFP4) mul_mat path with llama_prec_policy (llama/24364)
  • NVFP4: optimize MMQ accumulation and handle the compute type on the cuBLAS path (llama/29857, llama/29173)
  • Fuse shared experts into MMVQ (llama/29184)
  • Use MMVF for thin f16/bf16 mul_mat at small batch size (llama/29633)
  • FlashAttention: prefer whole-tile scheduling for two-stage kernels, tune fp16 tile configs for head sizes 40-112, fix 2 broken Volta cases (llama/29435, llama/26289, llama/29803)
  • Lightning indexer: halve the score memory (CUDA/Metal/Vulkan) and tile over keys and tokens for 4 heads (llama/29825, llama/29901)
  • Fuse RMS_NORM + SCALE into one kernel (llama/29393)
  • Bitonic argsort handles rows wider than one block (llama/28957)
  • Add F16 support to FWHT and a F16 kernel for CONV_2D_DW (llama/29096, llama/29064)
  • SSM_SCAN: support state size 96 (llama/28717)
  • Fixes: batch-independent alloc_deps check, unexpected graph reallocation in k-pool models, MMQ memory fault when n_expert >> n_ubatch, cpy transposed path corrupting non-contiguous dst (llama/29986, llama/29958, llama/29941, llama/27663)
  • Make the CCCL version configurable and pin it for CI (llama/29792)
  • HIP: enable the fattn-mma kernel on CDNA for dkq > 256 at large batch sizes, fixing template skip and DGX Spark misdetection (llama/28907, llama/29559, llama/29572)
  • HIP: optimize packed byte subtraction (llama/29478)
  • MUSA: fix PH1 (MTT S5000) operator and build issues, use the vector lightning indexer kernel (llama/29193, llama/29990)

Metal

Vulkan

  • int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (llama/27952)
  • Sparse flash attention for quantized K/V (llama/29639)
  • Perf: reuse descriptor sets when bindings are constant, fuse the SCALE鈫扴IGMOID鈫扴CALE鈫扝C_POST chain (llama/29280, llama/29520)
  • Flash attention: fix shmem write out of bounds and stale prealloc_y reuse across FA and soft_max (llama/29988, llama/29591)
  • Fix wrong results when a mul_mat reads a slice of a larger cache (llama/28956)
  • Tuning: GDN kernel on Intel, RDNA4 mat_vec, disable large matmul tile on Samsung GPUs with 32KB shared memory, argsort kernel selection on Adreno (llama/29476, llama/29934, llama/28531, llama/29469)
  • Fix build with legacy GLSLC without cooperativeMatrix API (llama/29409)

WebGPU

  • Add MMVQ support for Q1_0/Q5_0/Q5_1/Q3_K/Q5_K/Q6_K/MXFP4 (llama/29483)
  • Add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS and f16 support for fill/set_rows (llama/29358, llama/29897)
  • Fix SSM_SCAN binding aliasing and unaligned writes in buffer_set_tensor (llama/29750, llama/29471)

SYCL

  • Add sparse flash attention (llama/28796)
  • Q8_0 DMMV ESIMD and MMVQ wide load (llama/29186)
  • FWHT kernels for block widths above 512, large register file for D=512 FA vec kernels (llama/29243, llama/29062)
  • Avoid the slow oneDNN reference matmul and fattn (llama/28985)
  • Reduce tensor allreduce sync with pinned host buffers (llama/29604)

OpenCL

OpenVINO

  • Update backend to 2026.4.1: performance optimizations, expanded op support, improved device listing (llama/29852)
  • Serve GET_ROWS on a weight view from the base Constant (llama/28381)

Hexagon

RPC

  • Use the RDMA completion channel to avoid spinning, include nb in the get_alloc_size cache key and floor at ggml_nbytes (llama/29440, llama/29283)

zDNN

BLAS

  • Document AOCL-BLAS build and label the device AOCL-BLAS (llama/29640)

More info

Changelog since v0.25.3

d7cb5741 ggml : bump version to 0.26.0 (#1652)
77f2f491 sync : llama.cpp
316f7e78 CUDA: make the alloc_deps check batch independent (llama/29986)
9e974161 vulkan: fix Flash Attention shmem write out of bounds (llama/29988)
eb9fa483 vulkan: revert mul_mat_id tile selection PR #29182 (llama/29936)
ab367edb cuda: use the vector lightning indexer kernel on MUSA (llama/29990)
4c901461 CUDA: Optimize accumulation in mmq for NVFP4 type (llama/29857)
4bb5b76c vulkan: fix stale prealloc_y reuse across flash attention and soft_max (llama/29591)
4c526d30 vulkan: sparse flash attention for quantized K/V (llama/29639)
979ff228 llama : fix unexpected graph reallocation in the k-pool models (llama/29958)
28fb752c ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 (llama/28479)
f87a8749 vulkan : Fix undeclared identifiers when -DGGML_VULKAN_RUN_TESTS=ON (llama/29912)
d5486c5c webgpu: add MMVQ support for Q1_0/Q5_0/Q5_1/Q3_K/Q5_K/Q6_K/MXFP4 (llama/29483)
c26d651d cuda: tile the lightning indexer over keys and tokens for 4 heads (llama/29901)
e80a369b metal : few-row MMA mat-mul (llama/29869)
046df6f9 CUDA: use MMVF for thin f16/bf16 mul_mat at small batch size (llama/29633)
64b64bcc CUDA: prefer whole-tile FlashAttention scheduling for efficient two-stage kernels (llama/29435)
4ad8c550 CUDA: refactor swizzling code (llama/29612)
ce8ec243 sync : llama.cpp
ebefd1bb ggml-cpu: support BF16/FP16/FP32 K tails in tinyBLAS on x86 (llama/29806)
9dd7f1fb cuda : move neu_padded to where it is used (llama/29940)
799079d6 cuda : move blocks_per_col to where it is used (llama/29939)
51f4e809 CUDA: fix MMQ memory fault if n_expert >> n_ubatch (llama/29941)
8b304126 vulkan: fix rdna4 mat_vec tuning (llama/29934)
be35ddc7 webgpu: add f16 support to fill/set_rows (llama/29897)
51aab529 ci: fix flaky ADD_ADD f16 by using the fused ADD tolerance (llama/29904)
049c181a ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. (llama/29852)
84ae6f20 qwen4exp : halve the indexer score memory (llama/29825)
e45b9044 CUDA: fuse shared experts into MMVQ (llama/29184)
2ca61b70 ggml-cuda : fix cpy transposed path corrupting non-contiguous dst (llama/27663)
c30612b4 ggml-quants : avoid invalid rounding in qkx3 scale search (llama/29817)
8f987269 ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (llama/27096)
d981df5d metal : add tensor API flash attention kernel for F16 KV (llama/29570)
d1e4f33d opencl: use sigmoid f16 for bf16 (llama/29787)
fef80303 SYCL: Q8_0 DMMV ESIMD and MMVQ wide load (llama/29186)
6f7be931 vulkan: disable large matmul tile on Samsung GPUs with 32KB shared memory (llama/28531)
b420160a sycl: large register file for D=512 FA vec kernels (llama/29062)
55e936cc sycl : do not use slow oneDNN reference matmul and fattn (llama/28985)
5c05e590 qwen4exp : optimize mask constructions (llama/29824)
2264bd1c ggml : add alloc_buffer_n to buffer type interface (llama/23671)
826f7bfe vulkan: add logging to pipeline compile issues (llama/29794)
bf6fbd8a hexagon: install rebuilt HTP skels (llama/29828)
4e214c9f hexagon: add q2_k and q3_k quant type support (llama/29717)
fbaf749d CUDA: fix 2 broken Volta FA cases (llama/29803)
87c7bc20 llama: refer to segment documentation [no ci] (llama/29074)
57183ba2 hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (llama/29685)
5ad085c5 cuda : route sm70 to the Turing MMVQ nwarps table (llama/29753)
10ab2772 metal : release temporary private transfer buffers (llama/29777)
f3dc6fa4 webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- #29358 (llama/29358)
5f48c8fd CUDA: Handle compute type for NVFP4 on cublass path (llama/29173)
8e962d6c CUDA: Make CCCL configurable + pin it to 3.4.3 for CI jobs (llama/29792)
a5bac92a meta: clear inactive AllReduce shards with FILL, not SCALE (llama/29793)
e60a5a12 HIP: avoid treating CDNA as dgx spark for gqa_ratio 20 in fattn_mma dqk 576 (llama/29572)
ea00c4b0 hex-workqueue: fix race condition in seqn getting out of sync with idx_read/write (llama/29785)
d945687b BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS (llama/29640)
d5581dfa opencl: mark vec subgroup bcast as supproted for Adreno E17 compiler (llama/29698)
aea54a82 metal : use bf16 math for mxfp4 mul-mat (llama/29770)
b57d12f8 webgpu: fix SSM_SCAN binding aliasing (llama/29750)
ef74575b ggml-opencl : replace alloca() with std::vector (llama/29765)
410306d9 cuda: guard the iq4_nl dequantize row kernel against short rows (llama/29683)
c4fa3d2d Hexagon: optimize ALLREDUCE with support for safe scatter mode (llama/29757)
fa8f6fc9 ggml/gguf : fix integer overflow (llama/29384)
b5388253 ggml-et : remove useless alloca() (llama/29663)
bdda0794 ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) (llama/29675)
6bd4766f cpu: accept BF16 in src1 of mul_mat (llama/28937)
aeb2ec99 openvino: serve GET_ROWS on a weight view from the base Constant (llama/28381)
0c3606a3 musa : define CUDA_ARCH for device passes (llama/29508)
7724875d SYCL: reduce tensor allreduce sync with pinned host buffers (llama/29604)
3a9706a2 ggml-zdnn: impl buffer reset, fix memory leaks (llama/29637)
923daf2a Hexagon f16 activation ops (llama/29209)
566ef917 gguf : reject tensor size that wraps after padding (llama/26979)
3d3ab7b6 ggml : check row bounds in get_rows_back (llama/29575)
ee07a38d hexagon: optimize concat op (llama/29673)
ac1349f5 CUDA: bitonic argsort handles rows wider than one block (llama/28957)
52170272 ggml-cuda: HIP: optimize packed byte subtraction (__vsubss4 -> __vsub4) (llama/29478)
c3e8013e ci: add zdnn backend build but not test (llama/29541)
a9f73bc8 opencl: fix get_tensor for q5_K adreno gemm_nonshuffle kernel (llama/29555)
649077ea vulkan: Tune GDN kernel, fix Intel performance (llama/29476)
1b7dcfd5 vulkan : Load F32 A matrix 2 at a time when its 2-aligned (llama/29254)
f0877f7a vulkan: MOE aware mat_mul_id tile selection (llama/29182)
12c75a7d ggml : fix c++ odr by properly using GGML_COMMON_DECL_CPP (llama/29504)
29e99e83 ggml : accumulate f16 dot products in f32 on AVX512-FP16 (llama/29545)
3dcadb17 ggml : require input tensors to be GGML_OP_NONE (llama/29647)
92202620 hexagon: add FP32 GELU_ERF and GEGLU_ERF support (llama/29631)
f54e15b1 ggml : collect all input tensors into graph_inputs (llama/29634)
405955ee ggml-zdnn: fix 0-row tensor crash (llama/29636)
be1924eb ggml : speed up model loading (llama/29598)
33606a77 metal: FWHT perf optimizations (llama/29602)
a6a4a714 vulkan : reuse descriptor sets when bindings are constant (llama/29280)
64aaf933 vulkan: include functional header (llama/29597)
530645f0 ggml-openvino: mark unaligned batch-stride views unsupported (llama/29603)
88b69d4f webgpu: Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor (llama/29471)
741ad51d tests : refactor test-recurrent-state-rollback (llama/29426)
276472d0 ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 (llama/29423)
302722f9 vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain (llama/29520)
079c7a4a HIP: fix template skip for DKQ > 256 mfma kernels (llama/29559)
46fc5b3b metal: support left and circular padding in GGML_OP_PAD (llama/29561)
52d7fa07 Enables Windows ARM64 build with MSVC cl.exe (llama/28362)
05987856 vulkan: fix wrong results when a mul_mat reads a slice of a larger cache (llama/28956)
040a856e opencl: refine bin kernel loading condition (llama/29503)
da7ba41c sycl: FWHT kernels for block widths above 512 (llama/29243)
7f7f1b4f CUDA: tune fp16 tile FlashAttention configs for head sizes 40-112 (llama/26289)
94b651ec HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch sizes (llama/28907)
08a607df vulkan: fix argsort kernel selection for Adreno (llama/29469)
18a5df5c RPC: use RDMA completion channel to not spin (llama/29440)
3d6f297e hexagon: support tiled Q4_0 and Q8_0 GET_ROWS (llama/29511)
048577d2 hexagon: support for backend sampler (llama/29502)
2e573235 cuda: support Nemotron 3 Puzzle state size 96 for ssm scan (llama/28717)
c9719329 cuda: add F16 input to the FWHT (llama/29096)
37743976 ggml-cpu: tiled mul_mat for k-quants (llama/27851)
0e0d166c opencl: add A8 Q8_0 non-MoE dp4a binary kernel (llama/29439)
37f47b99 hexagon: find software divide calls using binary inspection tool (llama/29449)
e269548d opencl: add bin kernel kernel_gemm_noshuffle_q5_k_f32_32b_trans_ila_a8_bin, kernel_gemm_noshuffle_q5_k_q8_1_dp4a_ila_a8_bin (llama/29401)
40ccfe7c Fixing the vulkan build issue of legacy GLSLC version that has no cooperativeMatrix API support (ggml-org/llama.cpp#29373) (llama/29409)
a2904eec metal: FWHT kernels for block widths above 512 (llama/29095)
e81f673e metal : split fa kernels into per-dtype libraries (llama/29329)
d8fdaa85 llama : add llama_prec_policy + model-driven W4A4 path (llama/24364)
58bbebce HIP: bump HIP_VERSION requried for fp8 to avoid missing __hip_fp8_e4m3 support in 6.2 (llama/29231)
99bbbddd rpc: include nb in the get_alloc_size cache key and floor the result at ggml_nbytes (llama/29283)
ac0af6b5 support sparse FA (llama/28796)
aec2b25c musa: fix PH1 (MTT S5000) operator failures and build issues (llama/29193)
fa692ff7 CUDA: fuse RMS_NORM + SCALE into one kernel (llama/29393)
5fcf8056 hexagon: add q5_k quant type support (llama/29123)
89ff9882 hexagon: use DMA for contiguous dim1 CONCAT (llama/29404)
f11f23a8 metal : fix graph capture and handle empty graphs (llama/29390)
642b1d8f metal : optimize sparse FA + clean-up (llama/29377)
36a1d50e hexagon: handle multi-sequence in concat_2d (llama/29344)
ce159a17 hexagon: dynamic quantizer improvements (llama/29395)
2fef9fcd hexagon: support I32 CPY and CONT (llama/29379)
a99c77cf cuda : add F16 kernel support for CONV_2D_DW (llama/29064)
8885b063 test: flush status (llama/28352)
e2f1abba vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (llama/27952)

View the full release notes at https://github.com/ggml-org/ggml/releases/tag/v0.26.0.


@github-actions github-actions Bot added the bump-formula-pr PR was created using `brew bump-formula-pr` label Oct 5, 2026
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

馃 An automated task has requested bottles to be published to this PR.

Caution

Please do not push to this PR branch before the bottle commits have been pushed, as this results in a state that is difficult to recover from. If you need to resolve a merge conflict, please use a merge commit. Do not force-push to this PR branch.

@github-actions github-actions Bot added the CI-published-bottle-commits The commits for the built bottles have been pushed to the PR branch. label Oct 5, 2026
@BrewTestBot
BrewTestBot enabled auto-merge October 5, 2026 18:10
@BrewTestBot
BrewTestBot added this pull request to the merge queue Oct 5, 2026
Merged via the queue into main with commit 1b4e723 Oct 5, 2026
22 checks passed
@BrewTestBot
BrewTestBot deleted the bump-ggml-0.26.0 branch October 5, 2026 18:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bump-formula-pr PR was created using `brew bump-formula-pr` CI-published-bottle-commits The commits for the built bottles have been pushed to the PR branch.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants