Weekly GitHub Report for Llama.cpp: August 24, 2026 - August 31, 2026 (21:14:19)
Weekly GitHub Report for Llama.cpp
Thank you for subscribing to our weekly newsletter! Each week, we deliver a comprehensive summary of your GitHub project's latest activity right to your inbox, including an overview of your project's issues, pull requests, contributors, and commit activity.
Table of Contents
I. News
1.1 Recent Version Releases:
The current version of this repository is b4991
1.2 Version Information:
The version released on March 29, 2025, introduces key updates that enhance overall performance and user experience, reflecting a continued focus on stability and feature improvements. Notable highlights include optimized system processes and refined interface elements to streamline usability.
II. Issues
2.1 Top 5 Active Issues:
We consider active issues to be issues that that have been commented on most frequently within the last week. Bot comments are omitted.
-
[BUG-UNCONFIRMED] Eval bug: Memory Leak Issue Report for llama.cpp: This issue reports a persistent memory leak in llama-server where approximately 3MB of memory is leaked per request across multiple models and request types, causing out-of-memory (OOM) failures in long-running services. The leak is highly reproducible, affects both text and image requests, and results in continuous memory growth without reclamation, making the software unusable in production environments without frequent restarts or workarounds.
- The comments include verification of the leak through stress tests on Jetson Orin hardware, sharing of detailed logs and reproduction scripts, suggestions for environment variable tweaks and command-line options to mitigate the leak, and confirmation from multiple users experiencing similar OOM issues with different models and workloads.
- Number of comments this week: 9
-
[BUG-UNCONFIRMED] Eval bug: SYCL multi-GPU crash with Intel Arc Pro B50 + Arc A770: This issue reports a crash occurring when using the SYCL backend for multi-GPU inference with an Intel Arc Pro B50 and an Intel Arc A770, where single-GPU execution works correctly but multi-GPU execution fails with various errors depending on the split mode, including out-of-device-memory errors, segmentation faults, and backend assertions. The user seeks guidance on whether this multi-GPU configuration is expected to work with the current SYCL/Level Zero backend or if there are known limitations when combining these two different Intel GPU generations.
- The comments reveal that similar issues have been encountered by others, with some resolving crashes by rebooting or adjusting kernel drivers; suggestions include environment variable tweaks such as disabling host pinned memory and enabling virtual memory management, which led to a stable multi-GPU setup with improved throughput on a different model. Additional discussion covers benchmarking on dual B70 GPUs, configuration details, and confirmation that disabling host pinned memory is key to stability, while other SYCL features can remain enabled.
- Number of comments this week: 7
-
[BUG] llama-ui : unable to open reasoning level selection menu on desktop: This issue reports that after a recent update, the reasoning level selection menu in the llama-server's web UI is no longer accessible on desktop browsers, specifically macOS Safari, although it still appears on mobile devices. The problem seems to be a regression introduced by a specific commit, causing the reasoning menu to disappear when clicking the plus button, with the issue confirmed to be related to single-model mode.
- The comments include visual evidence of the missing menu, confirmation that the chat template supports reasoning options, observations that the menu appears on mobile and when resizing the desktop browser to a mobile width, and a recognition that this is a bug in single-model mode; a quick fix was proposed in a pull request addressing the regression.
- Number of comments this week: 7
-
[BUG-UNCONFIRMED] Vulkan GATED_DELTA_NET pipeline compile hangs on gfx1103 (RADV 780M) — llama-server never reaches listening: This issue describes a hang occurring when compiling the Vulkan GATED_DELTA_NET pipeline on an AMD Radeon 780M (gfx1103) GPU, where the llama-server process never reaches the listening state and pins a CPU core at 100% during model loading. The problem appears specific to certain large models and is not caused by the Vulkan driver version or quantization method, but rather seems related to how the GPU memory is allocated and managed during the loading process, with detailed debugging revealing GPU memory oscillations and command submission delays.
- The comments include offers to provide additional logs and debug information, suggestions to try newer drivers which did not resolve the issue, detailed follow-up investigations pinpointing the hang to GPU memory allocation and command submission in Vulkan, clarifications ruling out initial assumptions about the cause, and tests with environment variables that alter memory usage patterns but do not fix the hang, with ongoing requests for further targeted debugging.
- Number of comments this week: 7
-
[RESEARCH 🔬] Research: FreeToken is faster than llama.cpp: This issue discusses a research effort comparing the speed of FreeToken to llama.cpp, hypothesizing that FreeToken's availability renders llama.cpp obsolete due to its slower performance. The goal is to analyze FreeToken's methods to enhance llama.cpp's speed, with some progress made in background research and hypothesis formation but implementation and documentation still pending.
- The comments clarify that FreeToken is a Python-based project, which may limit direct relevance but offers techniques that could improve llama.cpp. Participants debate the claim that llama.cpp is deprecated, noting existing optimizations and ongoing work to implement similar improvements, while also discussing hardware requirements and potential use cases for FreeToken despite concerns about its current maintenance status.
- Number of comments this week: 6
2.2 Top 5 Stale Issues:
We consider stale issues to be issues that has had no activity within the last 30 days. The team should work together to get these issues resolved and closed as soon as possible.
As of our latest update, there are no stale issues for the project this week.
2.3 Open Issues
This section lists, groups, and then summarizes issues that were created within the last week in the repository.
Issues Opened This Week: 112
Summarized Issues:
- GPU Backend Crashes and Errors: Multiple issues report crashes, assertion failures, or fatal errors occurring in various GPU backends including HIP/ROCm, CUDA, Vulkan, SYCL, and Metal. These problems range from invalid argument errors, illegal memory accesses, kernel launch failures, to synchronization and race condition bugs, often triggered by specific hardware configurations, model parameters, or multi-GPU setups, severely impacting stability and performance.
- [issues/27670, issues/27674, issues/27678, issues/27698, issues/27717, issues/27749, issues/27750, issues/27757, issues/27761, issues/27763, issues/27769, issues/27771, issues/27792, issues/27829, issues/27833, issues/27835, issues/27849, issues/27901, issues/27905, issues/27910, issues/27911, issues/27918, issues/27923, issues/27927, issues/27931, issues/27888, issues/27953, issues/27964, issues/28018, issues/28047, issues/28048, issues/28056, issues/28060]
- Memory Management and Out-of-Memory Issues: Several reports highlight memory leaks, excessive VRAM or system RAM usage, and out-of-memory crashes across different platforms and backends. Problems include persistent memory leaks per request, VRAM overallocation causing model load failures, excessive disk writes on Windows due to mmap usage, and large memory reservations when using multiple GPUs or integrated plus discrete GPUs.
- [issues/27667, issues/27680, issues/27725, issues/27814, issues/27840, issues/27845, issues/28022, issues/28093]
- Speculative Decoding and MTP (Multi-Token Prediction) Bugs: Issues describe problems with speculative decoding implementations, including lack of speedup in server binaries, crashes with certain split modes, incorrect token acceptance past end-of-generation, and failures when models lack required MTP layers. These bugs affect generation correctness, performance, and server stability.
- [issues/27732, issues/27839, issues/27849, issues/28049, issues/28051, issues/28072]
- Model Output and Tokenization Errors: Multiple issues report degenerate or nonsensical model outputs, silent end-of-sequence token generation, infinite loops of repeated characters, and tokenization bugs such as stack overflows or control token warnings. These problems degrade user experience and model reliability, often linked to specific models or hardware.
- [issues/27683, issues/27733, issues/27756, issues/27876, issues/27772, issues/27797, issues/27767, issues/27783]
- Backend Performance Regressions and Optimization Research: Reports include performance regressions in CUDA buffer sizes, slower token generation on OpenVINO, under-parallelization on RDNA4 GPUs, and research efforts to improve flash attention on Volta GPUs and FreeToken speed advantages. These highlight ongoing challenges in optimizing backend throughput and efficiency.
- [issues/27682, issues/27680, issues/27796, issues/27837, issues/27685]
- Model Loading and Compatibility Issues: Several issues describe failures or crashes during model loading due to backend incompatibilities, driver mismatches, or unsupported model architectures. Problems include heap corruption on Windows Insider builds, unknown model architecture errors for GLM 5.3, and failures with specific MoE or Flash-Next models.
- [issues/27748, issues/27922, issues/27727, issues/27698, issues/27964]
- Server and RPC Stability Problems: Issues report server crashes, livelocks, and high CPU usage in idle states, including problems with RPC servers on Mac and WSL2, CUDA grid dimension overflows under concurrent load, and RPC device detection failures during system boot. These affect reliability in production deployments.
- [issues/27687, issues/27801, issues/27835, issues/27923, issues/28057, issues/28060]
- Tooling, Web UI, and API Bugs: Problems include failures in the WebUI copy-to-clipboard button, duplicate tool-call IDs in streamed responses, silent dropping of .webm video attachments, and bugs in grammar parsing or preset model deduplication. These issues impact usability and integration with external tools.
- [issues/27831, issues/27966, issues/28076, issues/28081, issues/27846, issues/27772]
- Build and Compilation Issues: Reports cover build failures due to missing CPU symbols when disabling CPU backend, compilation errors on ARM64 Apple M5 CPUs, and linking problems with backend DLL dependencies on Windows. These hinder cross-platform development and deployment.
- [issues/27834, issues/27982, issues/28009]
- Quantization and Model Conversion Bugs: Issues describe NaN logits during CUDA evaluation of quantized models, quadratic growth in quantization time on Windows, and prune-layers metadata mismatches causing model load failures. These affect model accuracy and conversion workflows.
- [issues/27899, issues/28034, issues/27976]
- Feature Requests and Enhancements: Requests include adding support for new models (Qwen3.8-Flash-Next, Tencent/WeMM-Embedding), CLI options to ignore sampling parameters, improved output formatting in performance tests, and generalizing file APIs for multimodal data. These aim to extend functionality and improve user experience.
- [issues/27741, issues/27938, issues/27893, issues/27904, issues/27942, issues/27987, issues/27989, issues/28051, issues/28090, issues/28087]
- Synchronization and Race Condition Bugs: Several issues describe race conditions causing cross-request contamination, deadlocks, or crashes, including missing synchronization in GPU backends, multi-GPU scheduling races, and recurrent state rollback corruption. These concurrency bugs lead to data corruption and instability.
- [issues/27750, issues/28019, issues/28056, issues/28047]
- Backend-Specific Hardware and Driver Issues: Reports include driver mismatches causing silent corruption, Vulkan driver hangs on AMD GPUs, and LoadLibraryW failing to find dependent DLLs on Windows due to search path limitations. These hardware and OS-specific problems complicate deployment.
- [issues/27763, issues/27998, issues/28009]
- Parsing and Grammar Bugs: Issues include parser functions incorrectly returning success on missing delimiters and grammar parsing failures causing fatal exceptions, leading to malformed output or HTTP 400 errors. These bugs affect input validation and robustness.
- [issues/27772, issues/28081]
- Speculative Decoding and Cache Management Bugs: Problems with ngram-cache not resetting between requests, causing stale data and reduced acceptance rates, and KV cache fragmentation causing throughput regressions on ROCm backend, degrade decoding quality and speed.
- [issues/27852, issues/27761]
- Miscellaneous Bugs and Investigations: Includes static analysis report issues, requests for prebuilt ROCm containers, and investigations into entropy-gated speculative decoding methods to improve efficiency.
- [issues/27740, issues/27821, issues/28087]
2.4 Closed Issues
This section lists, groups, and then summarizes issues that were closed within the last week in the repository. This section also links the associated pull requests if applicable.
Issues Closed This Week: 85
Summarized Issues:
- Model loading and compatibility issues: Several issues report problems with loading specific models or model formats, including unsupported model types causing server exits, crashes due to contradictory model metadata, and failures with new or hybrid model formats requiring explicit file specifications. These problems highlight challenges in model compatibility and the need for improved handling of diverse model architectures and formats.
- Backend-specific bugs and crashes: Multiple issues describe crashes and errors specific to various backends such as Vulkan, CUDA, SYCL, Metal, and ROCm, including memory access violations, driver or runtime errors, and resource allocation failures. These backend problems often cause system hangs, assertion failures, or incorrect computations, indicating instability and incompatibility in GPU and hardware acceleration layers.
- Performance regressions and optimizations: Several issues address performance problems such as slowdowns in prefill or decode phases, inefficient handling of large tool sets causing O(n²) complexity, and improvements like increasing workgroup sizes for Vulkan on specific GPUs. These highlight ongoing efforts to optimize inference speed and resource usage across different hardware and workloads.
- Cache and context management issues: Problems with caching and context checkpoint persistence are reported, including forced full prompt re-processing due to lack of cache data, failure to persist context checkpoints on slot restore, and loss of coherency in unified key-value caches during parallel runs. These issues cause performance degradation and incorrect model behavior during inference.
- Security vulnerabilities in ggml-rpc-server: Multiple security-related bugs are described, including null-pointer dereferences, use-after-free errors, out-of-bounds reads/writes, and heap-buffer-overflow vulnerabilities caused by improper input validation or unchecked parameters. These vulnerabilities can lead to server crashes or memory corruption when handling malicious RPC requests.
- Model inference and output correctness bugs: Issues include malformed output due to incorrect parsing of inline tags, degeneration to zero tokens in split RPC setups, and incorrect token selection caused by Vulkan graph optimizer bugs. These problems affect the quality and correctness of generated text and reasoning content.
- Tool and Web UI integration problems: Some issues report silent dropping of assistant messages containing tool calls, incorrect initialization of MCP/Tool toggles in the Web UI, and skipped UI settings for MCP servers in router mode, causing tool call failures and inconsistent user interface behavior.
- Build and compilation errors: Problems building llama.cpp with certain configurations are reported, including macro conflicts on Windows with BLAS support, missing HIP package files on Linux, and failures related to SYCL runtime libraries. These issues hinder successful compilation and deployment on various platforms.
- Antivirus false positives: Several Windows Defender and Bitdefender antivirus detections flag llama.cpp binaries as trojans, causing quarantine and installation issues despite indications that these are false positives confirmed by multiple users.
- Model and feature requests: Requests include adding support for new OCR models like PP-OCRv6 and MonkeyOCRv2, heuristic token drafting features for Gemma 4 MTP, and per-tensor mmap/read selection to improve loading performance and memory usage. These reflect community interest in expanding functionality and efficiency.
- Memory usage and leak concerns: Reports include large prompt cache entries causing expected but significant RSS growth, suspected memory leaks in Vulkan that were later attributed to caching behavior, and excessive default KV cache sizes in TTS models leading to high VRAM usage. These issues impact resource management and system stability.
- RPC and multi-node coordination errors: Crashes and failures occur in multi-node RPC setups due to tensor placement mismatches and backend CUDA graph compute errors during server interactions, indicating fragility in distributed inference configurations.
- Model evaluation and reranking issues: Problems with rerankers returning near-zero relevance scores due to missing tensor pooling, and evaluation crashes on specific hardware and software versions, highlight challenges in model evaluation pipelines.
- Miscellaneous bugs and errors: Additional issues include typos in Dockerfiles, backend column confusion in benchmarking tools, incorrect operation results on CPU and CUDA backends, and failures in vision support for specific models or file types. These diverse problems affect usability and correctness across the project.
2.5 Issue Discussion Insights
This section will analyze the tone and sentiment of discussions within this project's open and closed issues that occurred within the past week. It aims to identify potentially heated exchanges and to maintain a constructive project environment.
Based on our analysis, there are no instances of toxic discussions in the project's open or closed issues from the past week.
III. Pull Requests
3.1 Open Pull Requests
This section provides a summary of pull requests that were opened in the repository over the past week. The top three pull requests with the highest number of commits are highlighted as 'key' pull requests. Other pull requests are grouped based on similar characteristics for easier analysis. Up to 25 pull requests are displayed in this section, while any remaining pull requests beyond this limit are omitted for brevity.
Pull Requests Opened This Week: 126
Key Open Pull Requests
1. vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4: This pull request implements an int8 quantized cooperative matrix (coopmat1) matrix multiplication shader optimized specifically for AMD RDNA3 and RDNA4 architectures, significantly improving performance for various quantization formats and workloads such as Strix Halo and MoE prompt processing by hardcoding architecture-specific coopmat access patterns and enabling support for multiple quant types while selectively disabling slower configurations.
- URL: pull/27952
- Associated Commits: 13f6d, e086b, 5f034, 16256, 513f4, b121e, 2a878, 5553b, b453b, 14a60, 95d14, ad015, 31b9b, 23356, 9f069, 15b3c, 57b42, a9e0e, ef5c6, 2b9c5, 57dd3, ab6e8, 42a62, 18252, eb46b, d2a5a, ed294, fa37d, d8ed3, 3cd76, dc08f, ba840, 5037f, ab94f, 4695f, 22f58, 3985f, 09266, eeeac, 2ba1d, 9ef59, f2520, 0eab5, c30a8, 6421f, 04d7e, 584ae, e85ff, 8b194, 965e5
2. model: add GLM-5-Next (GLM-5.3-Flash): This pull request adds comprehensive support for the GLM-5-Next (GLM-5.3-Flash) model, a 321.3 billion parameter hybrid linear/sparse-attention Mixture of Experts (MoE) architecture including its vision tower, implements its specialized components such as the lightning indexer, KDA linear attention, MoE feed-forward layers with clamped SwiGLU activations, and dense DSA attention, integrates the model’s unique memory and caching mechanisms, updates the image preprocessor to match the reference, improves tokenizer handling for GLM-4, and includes extensive testing and performance optimizations to ensure correctness and efficient inference within the llama.cpp framework.
- URL: pull/27754
- Associated Commits: a505a, 9e1cc, 20a80, 1d99a, f320a, 7f256, cd69d, 6a583, 38415, 2db82, 83959, 8c498, fe959, 582fe, 9901a, 41cd9, 02fa4, 81e3e, 2d957, 6c59f, eab9e, 869e8, 282ef, cadbe, 204fa, ef531, e88c9, 2e0e5, cfa63, 6dd91, 1f0a3, b5517, f30be, 00699, d07e7, a175d, 57965, 949f7
3. Speculative prefill: This pull request implements Speculative Prefill, a method that accelerates long-context inference by using a smaller draft model to estimate token importance and selectively prefill only the most relevant context chunks into the main target model’s key-value cache, thereby significantly reducing compute time at the cost of some potential accuracy degradation.
- URL: pull/27692
- Associated Commits: a1f3a, 490ce, 061a3, ee7f3, 96be5, b3173, 4db32, dd7c7, 57b83, c4f08, e3c4a, 027ad, 2973a, cc935, 284a7, 29b0f, c0eda, 80a9f, ea8fc, 47c24, 0411a
Other Open Pull Requests
- Global system prompt loading in llama-server: This pull request adds the ability for the llama-server to load a global system prompt from a file using the
-sysfoption, allowing centralized control over system prompts across models. It also updates the WebUI to disable system prompt fields when this feature is used, while leaving non-chat endpoints unaffected.
- Build system improvements with precompiled headers: This update enhances the CMake build configurations by adding precompiled headers (PCH) for the most expensive headers in the Parsing/Frontend compiler stage to reduce compilation overhead. It also includes new scripts for building minimal and full versions of the project, excluding changes to the Codegen/Backend stage.
- New 32-element block quantization types: Two new 32-element block quantization types, IQ2_NL and IQ3_NL, are introduced for CPU, Metal, CUDA, and Vulkan backends to improve quantization efficiency on tensors with row lengths multiple of 32 but not 256. This lowers the performance floor caused by existing 32-block types and extends the non-linear quantization family downward.
- GLM-5.3-Flash (GLM5-Next) 320B hybrid model support: Support is added for the GLM-5.3-Flash 320-billion parameter hybrid model integrating text and vision capabilities by adapting existing llama.cpp components. The implementation includes new indexing and caching mechanisms, a Vision Tower with specialized preprocessing, and comprehensive testing to ensure compatibility and correctness.
- Qwen4Exp model correctness fixes: Multiple fixes improve the Qwen4Exp model implementation by enhancing sparse-attention block selection, supporting independent PLE embedding widths, validating metadata, updating caches, enabling recurrent-state rollback, rejecting unsupported tensor-split modes, and normalizing GDN queries and keys. These changes prevent inference errors, silent logit drift, and crashes while ensuring compatibility with official models.
- GLM-5.3-Flash model architecture additions: The GLM-5.3-Flash (
glm5next) model architecture is added, featuring KDA linear attention, MLA with DSA indexer and k-pool compression, gated-residual hyper-connections, and a 288-expert MoE. It reuses and extends existing subsystems to enable efficient pooled DSA selection, chunked token processing, and draft MTP head support, with thorough testing across hardware platforms.
- TQ1_0 ternary quantization support for Vulkan backend: Support for the TQ1_0 ternary quantization type is added to the Vulkan backend, including new data structures, shaders, pipeline integration, and test cases. This completes the implementation of this quantization format across all backends and improves decoding performance and accuracy on supported hardware.
- MTP draft head support for GLM-5.3-Flash model: This pull request adds support for the MTP draft head to the GLM-5.3-Flash model, enabling the
--spec-type draft-mtpoption and implementing efficient indexer reuse to improve performance. It demonstrates a 15-30% speedup with maintained draft acceptance rates on quantized layers.
- NUMA memory strategy "mirror" for ggml CPU backend: A new NUMA memory strategy called "mirror" is introduced that fully replicates large model weights on each NUMA node to enable simultaneous fast DMA-speed GPU prefill and NUMA-local CPU decode. This eliminates trade-offs and warmup delays of the existing
--numa distributeapproach, with automatic home-node selection and verified performance improvements on multi-socket GPU systems.
- Qwen3.8-Flash-Next model critical fixes: Multiple critical fixes address sequence copy indexing errors, block keying, image token position collapsing, exception handling for malformed metadata, and a CUDA abort caused by grid dimension limits. These fixes maintain perplexity and compatibility with existing GGUF files.
- Qwen4exp model generation performance optimizations: This pull request optimizes the qwen4exp model's long-context generation by reducing unnecessary full-cache scans, improving attention mechanisms, optimizing indexer summations, restricting predecessor scans, and replacing a std::set with a bitmap for used-cell tracking. These changes result in significant speedups in token generation and GPU utilization for very long contexts.
- Lumma-0.6B-Base multilingual model support: Support is added for the Lumma-0.6B-Base multilingual language model, incorporating Factorized Tied Embeddings and Shared KV Attention to optimize memory usage and deployment efficiency. Comprehensive testing ensures correctness against the reference HuggingFace implementation.
- AMD GCN-specific MMQ configuration for ggml-cuda HIP backend: A missing AMD GCN-specific MMQ configuration is added to properly handle wave64 execution and avoid fallback to less optimal RDNA2 MMQ config. Performance improvements include tile width adjustments and avoiding unnecessary fallbacks, resulting in significant matrix multiplication speed gains on AMD gfx906 devices.
- Fix for llama-server startup assertion with DFlash2 draft model: This fix ensures the lm_head is fully replicated on every device when using the DFlash2 draft model with
--split-mode tensoron CUDA, resolving assertion failures related toggml_top_k()on split outputs.
- Hugging Face Hub data layer for llama-ui: A new Hugging Face Hub data layer is added to
llama-ui, introducingHuggingFaceServicefor browsing and searching GGUF models with features like model details, repo file tree navigation, raw README fetching, and integration of the llama.app model catalog. It also includes GGUF file analysis helpers, new API types and enums, constants, and robust fetch retry logic.
- Intel iGPU memory crash and allocation fixes: This pull request fixes a crash caused by zero-size scratchpad memory and resolves the >4GB allocation limit on Intel iGPUs by ensuring valid memory pointers, replacing raw SYCL calls with managed ones, handling oneDNN virtual memory flags, and adding error handling to improve stability and compatibility on Intel UHD Graphics 770 iGPUs.
- Vulkan backend performance improvements on AMD GPUs: Performance is improved by copying strided f16 key-value data once into a scratch buffer to ensure sequential KV cache reads that engage more memory channels. This gating is applied specifically to AMD hardware based on measured benefits.
- Fixes for output_reorder() token indexing and lifecycle bugs: Critical indexing mismatches in
output_reorder()are fixed by introducing atoken_swapspermutation, snapshotting layout state to prevent mode-flip corruption, and clearing stale output permutations duringencode(). These changes prevent silent memory corruption and ensure consistent handling of output permutations.
- Transition of prebuilt UI hosting to GitHub Releases: The hosting of the prebuilt UI is moved from the Hugging Face bucket to GitHub Release artifacts by removing the
ui-publishjob, eliminating HF bucket-specific options, and updating the source build fallback to download checksum-verified UI tarballs from GitHub Releases with authentication. This simplifies the build process and improves artifact integrity verification.
- Qwen3 reranker memory allocation fix: The Qwen3 reranker is fixed to prevent exponential memory allocation growth with increasing physical batch size by removing unnecessary allocations, minimizing memory consumption without impacting normal LLM and embedding operations.
- Disabling K-quant optimization on Apple Silicon M3 Pro: The K-quant
mul_mv_extmatrix-vector multiplication optimization is disabled specifically for the Apple Silicon M3 Pro chip because benchmarking showed worse performance than the genericmul_mvpath. This improves performance by selecting the more efficient generic kernel for small batch sizes on M3 Pro devices.
- Standalone public C API for speculative decoding and MTP: A dependency-free public C API is introduced exposing multi-sequence speculative decoding and Multi-Token Prediction (MTP) capabilities extracted into
llama-speculative.cpp. This enables downstream projects to efficiently integrate draft-simple and MTP speculative decoding with native pointers, providing batching, state rollback, and verification for significant throughput speedups without bloating the core codebase.
- NextN/MTP draft head support for qwen4exp architecture: The NextN/MTP draft head with
--spec-type draft-mtpsupport is added for the qwen4exp architecture, tested on Qwen3.8-Flash-Next. It implements a new hyper-connection mixer and dense attention MoE block to enhance token prediction performance while maintaining compatibility with existing checkpoint formats.
- Tiled matrix multiplication for k-quant types: A tiled matrix multiplication implementation is introduced for k-quant types that unpacks quantized data in large 256x256 int8 tiles and applies an optimized 16x16 microkernel. This results in 3-7x CPU speedups on large matrices with minimal accuracy loss and portable C++ code adaptable to multiple architectures.
- Comprehensive Spark2_5 model support: End-to-end support for the Spark2_5 model is added, including GGUF conversion, architecture and tensor mappings, tokenizer integration, model loading, inference graph implementation, and extensive verification across CPU and CUDA platforms. This ensures parity and correctness without requiring new GGML operators or backend changes.
3.2 Closed Pull Requests
This section provides a summary of pull requests that were closed in the repository over the past week. The top three pull requests with the highest number of commits are highlighted as 'key' pull requests. Other pull requests are grouped based on similar characteristics for easier analysis. Up to 25 pull requests are displayed in this section, while any remaining pull requests beyond this limit are omitted for brevity.
Pull Requests Closed This Week: 182
Key Closed Pull Requests
1. spec-sidecar: improve RDNA2 DFlash and MTP sidecars: This pull request consolidates and improves the RDNA2 speculative-sidecar implementation by integrating HIP-based DFlash and MTP sidecars, adding a gfx1030 Q8_1/native-MMVQ DFlash path with split-K fallback, dynamically growing the DFlash KV cache beyond 16K positions, handling image and M-RoPE batches per sequence without permanently disabling sidecars, and keeping sidecars dormant unless explicitly enabled, all validated with performance benchmarks and correctness tests.
- URL: pull/28094
- Associated Commits: 8b24a, a313a, 3d23f, 400c4, 5d80b, f97f5, 3a2ab, 0b0c8, 81b07, 4936a, 1cd80, 6f711, 7390b, 35ee0, 667bc, 09f57, 10ce4, 2c8ba, 43427, 27f43, 7355f, 85833, ba3bf, a789f, fa1e9, 877a7, 0ec12, 45064, a9c80, 4b8ca, 337b8, 1b52c, 19373, d4cec, 3e286, 4c001, e4033, 61aae, 269c5, fec9d, a4f2d, a8378, 24b8c, 05a49, b2c53, 1e6a7, e2c43, f3387, 447e6, 457b3, 9bdd8, 4cb5e, 52ad2, 8f21e, f56e4, 6fb86, b6d36, 0711c, b4afc, ed356, 15834, c760f, 31f67, 22edd, 87365, 4e843, b4639, 3a8b3, 7dd6d, f8c7d, 45c24, 4139d, 4be2e, fae80, 4dd94, 1e5d0, 02206, ea6df, dad92, 60a1c, eccb5, 73c27, ee269, 180f1, f9b5b, 7805d, 19d26, a3478, e74f1, 29c56, b646b, 714a1, 0e6e8, c79b3, cba92, 7b36f, c1c3b, ff9b5, 6e0b6, f3de4, af27a, 39dda, 7ac36, db3eb, 22a7e, ae580, 2b020, b4f32, 0c78e, 970fd, adeef, 15df9, b1e44, 1d2c3, e4bc1, 75635, df252, 562c7, 0e082, aec7a, 1f321, c8870, ad7f5, 15937, 635bc, 91834, 7c5eb, 5abe1, 45a6d, eb580, fa7ee, d41b3, c5785, 1aa20, 83fe3, 5b0c0, ffb08, 2ab0e, 562ae, 9bbc5, f821c, e75bd, 8e566, 7b037, c7566, c8068, 414fc, f427d, dc256, bdce7, 99336, 5bcba, f7322, d2e97, 11524, e1702, c03a1, 86a42, e361e, 10557, cd740, f2a25, 97c80, 607ff, 922e1, ae782, 7989d, d58a0, 4f53e, c2da6, ccffd, 1eb7f, f9aaa, 77bdf, 91aa2, 8b668, 8d94a, 29be8, 26b42, b225c, 1faa0, 4caf0, 12b2f, a616d, a7c64, 6c26c, 98673, 5ad02, 01e75, e7b17, ccc3e, f07b6, 81d7d, 118ec, c9883, 357e7, 4f5f5, d31e2, 885fe, 73180, c543c, c7830, f97b4, 2a06e, ce231, 3fe3b, 9022a, 3aa11, d76e6, 8a864, 0e01d, d14c0, d986d, 6c8d8, b44a4, 4cd9c, 33dfa, 0bc9d, 64752, 57f10, 431dd, abcd7, 2b51a, 87133, 08988, 80122, 87662, 9a8d3, 661d1, 60b1d, 6499c, ed966, 209dd, c44f5, b3252, a4d7c, 56bcb, 50222, adc51, dfc73, 2bad4, 921a7, 3456f, 75874, bb9b8, 56539, f05d6, 0b54c, 1fd12, d4cd9
2. hexagon: support for multi-NPU devices (IQ9, IQ10) and fully asynchronous backend: This pull request overhauls the Hexagon backend to provide full support for multi-NPU devices such as the IQ9 and IQ10 series by implementing a fully asynchronous backend with async graph computation, events, tensor copies, cross-device synchronization, and a new device naming scheme, along with updated scripts and documentation to enable multi-NPU tensor-split mode and optimize performance and usability.
- URL: pull/26501
- Associated Commits: 6da87, 114d8, 40cfe, 65946, f181b, c1f66, edf22, 09e8d, 05667, db297, cae1f, 88dc3, 0800f, 3f4ed, f2a84, 9ad28, d5c5b, 70b96, 77035, 90c89, 46c75, 612ae, fe220, 234a9, 1b6e5, 1b70c, 57133, b6636, 1d008, cbd7a, a2a5c, 4a95c, 922ef, 68c43, 3f842, d593a, a7e83, e025e, ca7e0, cda0a, 3cdaa, d57d0, 4bad4, 4b14b, fc212, 265e6, 4af4f, 6976c, 1ab0d, c838c, dacc0, 1132c, 15b7b, 868d1, 4ddd8, f7b49, 05b71, 24ef2, ed0d3, 10d87, e629a, 37aee, dc2e2, df14a, 45ff9, 2e418, 0a2de, e216b, f9f95, ab41f, e8bae, 91525, 55f0b, 71851, b1882, 0288c, 673c9, 2a11b, 9fdca, 0afb0, 7bc85, 490de, db719, 4f4a5, 44cdd, c3341, 57313, 71854, 7156a, 9aa6b, 5efda, 029ce, c6f81, 03c69, d9fc4, 0ee17, c00ac, c4b5b, e02f6, 11089, f8d57, d640a, 11433, 1b832, a6eb4, 41985, d8fa0, 57749, dd756, 55993, 85ec5, a50e1, 50d46, 34b08, 3fc24, 62e05, f0d18, 8aa79, da43b, 891eb, e331d, 3b5b7, e696a, 40da2, 46123, 1544b, dfd01, a62ac, 4c708, 750bd, 5ea5c, 2aa7b, 976eb, f5c2f, 5adb0, bc434
3. ignore: This pull request introduces and ports AMD-native ROCmFPx FPx tensor quantization types with CPU reference codecs and Vulkan kernels into the llama.cpp project, along with multiple Vulkan performance optimizations, bug fixes, and benchmarking infrastructure, significantly improving prefill throughput and supporting advanced speculative decoding features while maintaining compatibility with upstream changes.
- URL: pull/28052
- Associated Commits: 64c3b, 8610c, a0374, e282d, aedc9, 82ac8, 3b804, f0a2b, 5a402, 9298f, c2e32, eacc5, c96f8, 2cef9, 25748, 5a108, 8aeac, f8982, d07d1, 4801d, e3c87, 16f07, ca261, 0b120, 22ed5, f97c0, 1674a, e98e6, 20c9a, eb3d3, aeca2, 68fe4, be760, e084d, 88ed6, bc533, a8c89, 9dfbd, 3bddc, e1cac, 28a3e, 54ca3, 2129b, 80c2c, 67483, 27ee6, 5a09c, c28d5
Other Closed Pull Requests
3.3 Pull Request Discussion Insights
This section will analyze the tone and sentiment of discussions within this project's open and closed pull requests that occurred within the past week. It aims to identify potentially heated exchanges and to maintain a constructive project environment.
Based on our analysis, there are no instances of toxic discussions in the project's open or closed pull requests from the past week.
IV. Contributors
4.1 Contributors
Active Contributors:
We consider an active contributor in this project to be any contributor who has made at least 1 commit, opened at least 1 issue, created at least 1 pull request, or made more than 2 comments in the last month.
If there are more than 10 active contributors, the list is truncated to the top 10 based on contribution metrics for better clarity.
| Contributor | Commits | Pull Requests | Issues | Comments |
|---|---|---|---|---|
| 477 | 0 | 0 | 0 | |
| allozaur | 328 | 27 | 0 | 39 |
| ngxson | 279 | 11 | 0 | 17 |
| ggerganov | 179 | 13 | 0 | 64 |
| edwinbrowwn | 219 | 1 | 0 | 0 |
| max-krasnyansky | 142 | 3 | 0 | 1 |
| danielhanchen | 111 | 5 | 0 | 0 |
| Summer110622 | 102 | 0 | 0 | 0 |
| ServeurpersoCom | 70 | 8 | 0 | 7 |
| 0cc4m | 60 | 3 | 0 | 18 |