forked from ggml-org/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 8
Pull requests: AMD-Ecosystem/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
ci: re-enable ubuntu-22-rocm release build with larger ccache
#91
opened Aug 18, 2026 by
jimw567
Collaborator
Loading…
3 tasks
cuda/hip: optional wave64 for the quantized mat-vec kernels on RDNA
#89
opened Aug 17, 2026 by
mgehre-amd
Collaborator
•
Draft
ggml-cuda: widen Q4_K MMVQ weight loads to global_load_b128
#88
opened Aug 17, 2026 by
Annieren
Loading…
CUDA: dispatch the chunked gated_delta_net in CU mode when it fills the device
#79
opened Jul 30, 2026 by
roberteg16
Loading…
llama: let the backend pick n_ubatch, and default to 2048 on RDNA3.5
#77
opened Jul 29, 2026 by
roberteg16
•
Draft
3 tasks done
tools: add mmq-tune, an MMQ tile-width autotuning harness
#76
opened Jul 29, 2026 by
roberteg16
•
Draft
3 tasks done
ggml-cuda: f32 tall-skinny GEMM kernel for gfx1151 (RDNA3.5)
#74
opened Jul 28, 2026 by
roberteg16
•
Draft
feat(cuda): fuse activations and residual add into mmv f/q epilogues
#67
opened Jul 24, 2026 by
roberteg16
•
Draft
tests: add MoE MMQ benchmark with routing-distribution generator
#62
opened Jul 20, 2026 by
roberteg16
•
Draft
3 tasks
gfx1151 (Strix Halo): fuse attn_k+v into single MMVQ dispatch
#59
opened Jul 17, 2026 by
jeffli-xilinx
Loading…
7 tasks done
ggml-cuda: GEMM weight row padding + one-time K-padded f16 dequant for prefill
#57
opened Jul 17, 2026 by
roberteg16
•
Draft
4 of 5 tasks
ProTip!
Exclude everything labeled
bug with -label:bug.