Weekly Project News

Archives
Subscribe

Weekly GitHub Report for Llama.cpp: September 21, 2026 - September 28, 2026 (20:38:35)

Weekly GitHub Report for Llama.cpp

Thank you for subscribing to our weekly newsletter! Each week, we deliver a comprehensive summary of your GitHub project's latest activity right to your inbox, including an overview of your project's issues, pull requests, contributors, and commit activity.


Table of Contents

  • I. News
    • 1.1. Recent Version Releases
    • 1.2. Other Noteworthy Updates
  • II. Issues
    • 2.1. Top 5 Active Issues
    • 2.2. Top 5 Stale Issues
    • 2.3. Open Issues
    • 2.4. Closed Issues
    • 2.5. Issue Discussion Insights
  • III. Pull Requests
    • 3.1. Open Pull Requests
    • 3.2. Closed Pull Requests
    • 3.3. Pull Request Discussion Insights
  • IV. Contributors
    • 4.1. Contributors

I. News

1.1 Recent Version Releases:

The current version of this repository is b4991

1.2 Version Information:

The version released on March 29, 2025, introduces key updates and improvements, focusing on enhanced functionality and performance optimizations. Notable highlights include streamlined features aimed at improving user experience and system efficiency.

II. Issues

2.1 Top 5 Active Issues:

We consider active issues to be issues that that have been commented on most frequently within the last week. Bot comments are omitted.

  1. [BUG-UNCONFIRMED] eval bug: draft-mtp roughly halves prompt processing on multi-GPU layer split (single GPU is fine): This issue reports a significant performance regression when using the draft-mtp speculative decoding mode with multi-GPU layer splitting, where prompt processing throughput roughly halves compared to no MTP or single-GPU setups. The user provides detailed benchmarks, analysis, and hypotheses that the synchronous catch-up decode in the MTP process breaks pipeline parallelism across GPUs, causing large overheads, and shares attempts at fixes and profiling that isolate the problem to synchronization and scheduling inefficiencies in multi-GPU contexts.

    • The comment discussion extensively investigates the root cause of the slowdown, confirming that the per-ubatch catch-up decode is not the dominant cost but that the target model’s own decode call takes significantly longer with MTP on multi-GPU splits. Proposed fixes involve lagging the MTP catch-up decode by one batch to avoid pipeline stalls, excluding the MTP context from pipeline parallelism, and improving graph reuse, which yield substantial prompt processing speedups without affecting generation quality. Several users test these patches with mixed results depending on model architecture and hardware, noting that increasing logical batch size can partially mitigate the issue on master, but some builds and configurations still suffer from degraded performance or instability. The conversation highlights the complexity of multi-GPU scheduling with MTP and the need for maintainers’ input to finalize robust fixes.
    • Number of comments this week: 20
  2. [ENHANCEMENT] Feature Request: Add support for K2 Horizon (0.9B, 3.7B, 7B, 32B, 36B MoVA): This issue requests the addition of support for the K2 Horizon model family, which includes five sizes ranging from 0.9B to 36B parameters, into the llama.cpp project to avoid maintaining a separate fork and ensure compatibility with upstream updates. The feature involves implementing a new architecture with specific modifications such as grouped RMSNorm, an optional softplus gate on attention output, and a novel MoVA (mixture of value experts) mechanism for the largest model, with all necessary code and tokenizer changes already developed and tested on CPU and CUDA.

    • The comments focus on clarifying the reasons for modifying the chat template due to unsupported templating features in the original llama.cpp, discussing fixes for sameas and dict() support, and agreeing on submitting a single PR with suggested code formatting improvements before merging.
    • Number of comments this week: 10
  3. Qwen3.8-27B (hybrid Gated DeltaNet): decode throughput collapses ~25x at context positions >~80K while prompt processing stays fast: This issue reports a severe and nonlinear collapse in decode throughput for the Qwen3.8-27B hybrid Gated DeltaNet model on CUDA-enabled RTX 4080 SUPER GPUs when decoding beyond approximately 80K context tokens, while prompt processing remains fast. The problem appears positional rather than memory-related, with extensive testing ruling out quantization levels and Flash Attention flags as causes, and subsequent investigation suggests the original cliff may have been due to environmental or measurement factors rather than a fundamental llama.cpp bug.

    • Commenters attempted reproductions across various hardware, drivers, and configurations, with some initially suspecting an NVIDIA driver regression but later retracting that after correcting timing metrics; multiple independent tests on CUDA, Vulkan, and HIP backends failed to reproduce the severe decode cliff, instead showing smooth throughput degradation or stable performance, and some noted a CUDA-specific penalty for quantized KV formats but no catastrophic collapse, leading to a consensus that the original issue is likely environment-specific or resolved.
    • Number of comments this week: 9
  4. Compile bug: Vulkan fails with a glslc lacking GL_KHR_cooperative_matrix since #24406 (fa_decode shaders built as coopmat unconditionally): This issue reports a compile bug in the Vulkan backend where shaders are unconditionally built with cooperative matrix support despite the GLSLC compiler lacking the necessary GL_KHR_cooperative_matrix extension, causing build failures on toolchains with older GLSLC versions. The problem stems from missing preprocessor guards around certain shader compilations and pipeline creations introduced in a recent commit, which breaks compatibility with environments like the Android NDK's bundled GLSLC.

    • The comments discuss acknowledging the issue, referencing a fix in a related pull request, requests for a separate quick-merge PR, inquiries about testing with older Vulkan SDK versions, and verification efforts to ensure the fix properly addresses the missing preprocessor guards.
    • Number of comments this week: 6
  5. [BUG-UNCONFIRMED] Misc. bug: Openvino cannot run gemma on intel core 7 155h: This issue reports a limitation in the OpenVINO NPU plugin on the Intel Core Ultra 7 155H, where the device supports only up to 255 command queues per process, causing failures when running the gemma model that requires compiling over 700 models simultaneously. The user explains that this queue limit leads to misleading error messages and forces fallback to CPU execution, and requests clarification on whether this limit is hardware-imposed or configurable, as well as better error reporting and documentation.

    • The comments discuss related issues with running advanced models on the NPU, including hangs and device resets with GatedDeltaNet models, difficulties with outdated transformer support in OpenVINO, and potential workarounds such as direct NPU programming or alternative backends; overall, the conversation highlights the current hardware and software constraints limiting effective NPU use for large language models.
    • Number of comments this week: 4

2.2 Top 5 Stale Issues:

We consider stale issues to be issues that has had no activity within the last 30 days. The team should work together to get these issues resolved and closed as soon as possible.

As of our latest update, there are no stale issues for the project this week.

2.3 Open Issues

This section lists, groups, and then summarizes issues that were created within the last week in the repository.

Issues Opened This Week: 79

Summarized Issues:

  • Multi-GPU and CUDA Backend Stability Issues: Several issues report crashes, assertion failures, and performance degradations related to multi-GPU setups and CUDA backends. Problems include invalid argument errors during kernel launches, deadlocks, and crashes during prefill or decode phases, often linked to tensor splits, batch shape limits, or GPU driver resets.
    • issues/29235, issues/29255, issues/29418, issues/29466, issues/29562
  • Backend-Specific Bugs and Crashes (SYCL, Vulkan, Metal): Multiple backend implementations exhibit crashes and incorrect behavior, including Vulkan memory reporting errors, SYCL backend crashes on convolution and speculative decoding, Metal backend memory mapping issues, and Vulkan fence timeouts causing decode failures. These bugs often cause crashes, memory errors, or degraded performance on specific hardware or driver versions.
    • issues/29241, issues/29277, issues/29352, issues/29366, issues/29373, issues/29452, issues/29455, issues/29465, issues/29526, issues/29527
  • Model Parsing and Reasoning Logic Failures: Issues highlight problems with tool call parsing in chain-of-thought reasoning, inconsistent enforcement of tool_choice requirements, and requests for UI commands to control thinking levels. These affect agentic workloads and user interaction with reasoning-capable models, causing failures or usability challenges.
    • issues/29248, issues/29262, issues/29264, issues/29295
  • Memory Management and Allocation Bugs: Several bugs involve incorrect memory allocation or management, such as stale allocation plans causing output corruption, excessive memory usage due to unbounded buffers or incorrect cache limits, and mmap mapping errors leading to GPU memory exhaustion. These issues cause crashes, corrupted outputs, or inefficient resource use.
    • issues/29313, issues/29322, issues/29324, issues/29388, issues/29391, issues/29465, [issues/29487](https://github.com/issues/29487], issues/29494
  • Performance Regressions and Optimization Requests: Reports include throughput regressions in CUDA and Vulkan builds, inefficiencies in flash attention with quantized caches, and requests for hardware-specific performance improvements such as Vulkan coopmat on Intel processors and CPU Flash Attention updates. These affect speed and resource utilization during inference.
    • issues/29251, issues/29326, issues/29341, issues/29342, issues/29371, issues/29378, issues/29310, issues/29349, issues/29410
  • Server and API Usability and Stability Issues: Several issues describe server crashes, hangs, or denial-of-service conditions caused by unbounded token counts, grammar parsing inefficiencies, improper request handling, and header misuse. There are also feature requests for better metrics, UI improvements, and command additions to enhance user experience and server robustness.
    • issues/29318, issues/29323, issues/29372, issues/29451, issues/29456, issues/29457, issues/29458, issues/29462, issues/29495, issues/29498, issues/29499, issues/29505, issues/29581
  • Model Loading and Initialization Failures: Issues include crashes or slowdowns during model loading or first inference calls, out-of-memory errors due to improper memory fitting, and invalid tensor parsing causing runtime errors. These problems affect startup stability and initial performance of models on various platforms.
    • issues/29261, issues/29345, issues/29383, issues/29438, issues/29513, issues/29521, issues/29579
  • Quantization and Numeric Precision Issues: Bugs include precision loss in GELU activation due to half-precision lookup tables, sign errors in int8 dot product implementations, and numeric noise in vision embedding paths causing accuracy drops. These affect model correctness and output quality.
    • issues/29251, issues/29252, issues/29351
  • Feature Requests for Model and Hardware Support: Requests include adding support for new model families like K2 Horizon, Qualcomm NPU integration, and non-uniformly compressed models with mixed precision quantization, aiming to expand the project's capabilities and hardware compatibility.
    • issues/29424, issues/29436, issues/29441
  • Build and Installation Issues: Problems include missing header installations due to CMake misconfiguration and compile failures related to backend shader builds or compiler/runtime bugs, impacting developer experience and build reproducibility.
    • issues/29350, issues/29366, issues/29373
  • Crash and Deadlock on Specific Hardware or Configurations: Reports include crashes on Apple A15 iOS devices, deadlocks on Intel Arc GPUs with SYCL/Level Zero, and reboot issues on NVIDIA Jetson Orin NX, often requiring workarounds or service restarts to maintain operation.
    • issues/29438, issues/29527, issues/29499, issues/29549
  • Sampling and Token Generation Bugs: Issues include repeated token outputs, incorrect end-of-generation token handling, and excessive memory allocation for top-k buffers, causing incorrect or inefficient sampling behavior.
    • issues/29392, issues/29577, issues/29487
  • Command Line and Completion Enhancements: Requests for bash-completion improvements and new commands to control reasoning levels and thinking tags aim to improve developer and user interaction with the software.
    • issues/29289, issues/29262, issues/29264
  • Miscellaneous Bugs and Issues: Various other bugs include broken documentation links, sign errors in rounding operations, and silent server terminations due to assertion failures or memory errors, affecting usability and stability.
    • issues/29505, issues/29406, issues/29391, issues/29322

2.4 Closed Issues

This section lists, groups, and then summarizes issues that were closed within the last week in the repository. This section also links the associated pull requests if applicable.

Issues Closed This Week: 104

Summarized Issues:

  • ROCm and HIP GPU Crashes and Compatibility Issues: Multiple issues report crashes and errors when running models on AMD GPUs using ROCm and HIP backends, including segmentation faults with multi-GPU setups, assertion failures during model loading, and kernel incompatibilities causing crashes or degraded performance. These problems often relate to backend support for specific hardware, driver versions, or multi-GPU memory management, requiring workarounds like disabling features or reverting commits.
    • issues/15845, issues/17583, issues/20433, issues/24132, issues/25807, issues/26435, issues/29552
  • Model Loading and Memory Allocation Failures: Several issues describe crashes, assertion failures, or errors during model loading or memory allocation across various backends including Vulkan, OpenVINO, and CUDA, often triggered by unsupported features like multi-buffer allocations, shape inference errors, or memory fitting probes. These failures prevent models from running correctly and sometimes require disabling certain features or backend fallbacks.
    • issues/22197, issues/24343, issues/24415, issues/25777, issues/29259, issues/29270
  • Multi-GPU and GPU Identification Problems: Issues highlight segmentation faults, crashes, or performance problems when using multiple GPUs, including incorrect GPU classification (e.g., NVIDIA Blackwell GPUs misidentified as integrated), causing failures in tensor split modes or memory allocation. These problems affect model loading and inference stability on multi-GPU systems.
    • issues/17583, issues/25751, issues/26901, issues/29260
  • Backend-Specific Performance Regressions and Inefficiencies: Various backends such as SYCL, Vulkan, CUDA, and ARM show performance regressions or inefficiencies, including high GPU power states, slow decode throughput, and increased VRAM usage during prompt processing. Some regressions are linked to recent commits or driver interactions, with partial workarounds like disabling features or adjusting flags improving throughput.
    • issues/24946, issues/26435, issues/26747, issues/26795, issues/27086, issues/29335
  • Speculative Decoding and Draft Model Bugs: Several issues report crashes, acceptance rate drops, or load failures related to speculative decoding implementations and draft models, including problems with tokenizer mask token IDs, invalid vector subscripts, and missing context leading to crashes or degraded speculative decoding behavior.
    • issues/26475, issues/26478, issues/26761, issues/26884, issues/26913, issues/27089
  • KV Cache Handling and Quantization Bugs: Issues describe bugs in key-value cache compression, quantization, and state restoration, including silent desynchronization, incorrect GPU usage during compression, and missing support for new quantization formats, which cause inference errors, crashes, or degraded accuracy during long-context inference.
    • issues/19466, issues/26402, issues/26777, issues/26989
  • API and Feature Requests for Improved Compatibility and Monitoring: Requests include adding native support for OpenAI Responses API, exposing detailed per-device memory usage metrics, adding JSON arguments for multimodal prompt concatenation, and Prometheus metrics for request context sizes, aiming to improve integration, observability, and prompt handling in llama.cpp server.
    • issues/19138, issues/26008, issues/26129, issues/27011
  • Router Mode and Model Preset Configuration Issues: Problems include router mode failing to load models from presets, regression in model preset naming causing broken routing, and draft model deduplication bugs, leading to configuration mismatches and failed model loading in server environments.
    • issues/25150, issues/27846, issues/29225
  • Vulkan Backend Specific Bugs and Crashes: Multiple Vulkan backend issues cause crashes or incorrect outputs due to unsupported tensor view offsets, stale cache entries affecting rollback, and numerical errors in dequantization paths, resulting in assertion failures, corrupted outputs, or server aborts.
    • issues/26685, issues/26744, issues/26853, issues/27022
  • Server Stability and Crash Issues: Various crashes occur due to memory access faults, assertion failures, heap corruption, and race conditions in server components, including crashes on startup, during inference, or on exit, often requiring restarts or disabling features to maintain stability.
    • issues/26782, issues/26906, issues/27065, issues/29138, issues/29188, issues/29552
  • Build and Compilation Failures: Several issues report build failures on specific platforms or backends, including Windows ARM64 assembly macro issues, CUDA backend compilation errors for sm_70 architecture, and HIP backend incompatibilities with older ROCm versions, blocking successful builds without patches.
    • issues/26875, issues/29222, issues/29223, issues/27303
  • Tokenization and Parsing Bugs: Bugs include tokenizer stack overflows due to excessive recursion, incorrect parsing of unary minus in templates, unsigned integer truncation in grammar parsing, and special tokens being misinterpreted as actual tokens, causing crashes, incorrect outputs, or prompt injection vulnerabilities.
    • issues/26965, issues/29233, issues/29380, issues/29546
  • UI and User Experience Issues: Problems include inability to edit multi-line prompts in terminal, silent waiting when starting a second server instance on an occupied port, UI MCP call timeouts with many servers, and missing SVG elements in downloaded images, causing user confusion or degraded interface functionality.
    • issues/26822, issues/26935, issues/26999, issues/28336
  • Security Vulnerabilities in RPC Backend: Out-of-bounds read and write vulnerabilities in ggml RPC backend operations allow potential memory corruption or data leakage due to missing runtime checks on tensor shape compatibility in GRAPH_COMPUTE requests, posing security risks.
    • issues/26825, issues/26912
  • Miscellaneous Bugs and Requests: Other issues include bugs in adaptive P setting not affecting output, inability to input mp4 files, missing Windows Arm Hexagon NPU official releases, and requests for improved model modularization and SSE4.1 optimizations for older CPUs.
    • issues/26776, issues/26752, issues/26877, issues/28772, issues/27066, issues/27055

2.5 Issue Discussion Insights

This section will analyze the tone and sentiment of discussions within this project's open and closed issues that occurred within the past week. It aims to identify potentially heated exchanges and to maintain a constructive project environment.

Based on our analysis, there are no instances of toxic discussions in the project's open or closed issues from the past week.


III. Pull Requests

3.1 Open Pull Requests

This section provides a summary of pull requests that were opened in the repository over the past week. The top three pull requests with the highest number of commits are highlighted as 'key' pull requests. Other pull requests are grouped based on similar characteristics for easier analysis. Up to 25 pull requests are displayed in this section, while any remaining pull requests beyond this limit are omitted for brevity.

Pull Requests Opened This Week: 121

Key Open Pull Requests

1. chat : implement token aware parsing: This pull request implements token-aware parsing for chat outputs by maintaining a token map alongside generated text to accurately distinguish special tokens from regular text, replaces regex-based trigger pattern matching with a more efficient token-based scanning grammar, refactors parsing utilities for improved ergonomics, and migrates existing parsers to leverage these token-aware mechanisms.

  • URL: pull/29307
  • Associated Commits: 61c94, 7e10e, e5e28, 81de3, faa17, 7aa96, 65c64, 5e450, cade2, 06a9d, f051a, 6b705, bf8f8, 1c4e8, 59940, 40630, 6ccbd, c2399, f3942, 05149, a5d81, a4fcf, b688d, d91c4, bed67, 3810f, 20e9d, e5457, 17105, 24bd0, 9ee70, cfb18, 483e6, 777ba, 191a4, 707b2, 1107a, ec7ad, e8437, 7cf9a, 3a10b, 3bc19, 48147, 305e4, 2f910

2. model: GraniteSpeech5ForCTC (Turbo CTC): This pull request adds support for the non-autoregressive, encoder-only GraniteSpeech5ForCTC model architecture to the llama.cpp project, implementing new code paths for CTC-based speech-to-text transcription including a dedicated server route, specialized decoding methods, and necessary adjustments to batching and model output handling to accommodate this fast, logits-based model that bypasses traditional LLM backbones.

  • URL: pull/29446
  • Associated Commits: a6c15, 8ba96, 899c7, e16bb, 80dca, 77d51, 2db40, 169fc, 59d9f, bd382, f96a8, 22929, a8cd8, b4899, 947eb, 21925, bf397, a5f13, b10fd, d9335, ad5b5, e28fb, 24e13, ede31, 53a47, ab1cd, 11b9a, 486e6, 0fa21, 9761d

3. SVE Implementation of gemv q6 k 8x8 q8 k: This pull request implements support for SVE 128 and SVE 256 vector extensions in the gemv_q6_K_8x8_q8_K kernel, demonstrating performance improvements over NEON on various thread counts and hardware configurations while maintaining equivalent perplexity.

  • URL: pull/29256
  • Associated Commits: 70e7e, e601f, f04ab, ff5ae, d1224, 97ff3, abcd5, 1eb10, acb88, a86f5, 575ea, 1dda3, e9616, cb394, f732c, 0a7b4, e54b0, 0af25, e5367, a4ca4, fbc84

Other Open Pull Requests

  • Code Refactoring and Modularization: Multiple pull requests focus on improving code maintainability and organization by refactoring large files into smaller components and introducing common interfaces. These changes enhance navigation and simplify maintenance without altering public APIs or core functionalities.
    • pull/29354, pull/29385
  • Model Support and Integration: Several pull requests add support for new models such as K2 Horizon, MiniCPM v4.7, Laya multilingual decision model, and Limite 1B Violetto. These include model converters, tokenizer integration, inference runtimes, and template updates to ensure compatibility and correct functionality.
    • pull/29535, pull/29416, pull/29359, pull/29433
  • Performance Improvements and Optimizations: Multiple pull requests introduce performance enhancements such as improved CUDA kernels, multi-token decode optimizations, memory usage reductions, and grouped GEMM kernels for MoE models. These optimizations result in faster token processing, reduced memory consumption, and better hardware utilization.
    • pull/29353, pull/29367, pull/29578, pull/29470
  • Continuous Integration and Build Enhancements: Updates to the CI pipeline include adding IBM zDNN backend builds, upgrading oneAPI Toolkit versions, fixing compiler warnings, and addressing SYCL build issues. These changes ensure compatibility with new hardware and software environments while maintaining build stability.
    • pull/29541, pull/29273, pull/29363, pull/29541
  • RPC and Server Improvements: Enhancements to the RPC system and server include caching multiple graphs by UID to improve decoding speed, faster RPC loading times, prompt prefix preloading, and better handling of unmatched media markers in chat text. These changes improve responsiveness and robustness of server operations.
    • pull/29287, pull/29413, pull/29578, pull/29247
  • Bug Fixes and Security Hardening: Several pull requests fix critical bugs such as integer overflow vulnerabilities, null pointer dereferences, incorrect int8 handling, and security issues in audio projector files. These fixes improve stability, correctness, and security of the codebase.
    • pull/29359, pull/29421, pull/29360, pull/29291, pull/29421
  • New Features and Interfaces: Introduction of new tools and interfaces includes a typed decision readout CLI for Jev-like models and mechanisms for caching and managing prompt prefixes. These features expand the framework's capabilities without impacting existing APIs.
    • pull/29321, pull/29578

3.2 Closed Pull Requests

This section provides a summary of pull requests that were closed in the repository over the past week. The top three pull requests with the highest number of commits are highlighted as 'key' pull requests. Other pull requests are grouped based on similar characteristics for easier analysis. Up to 25 pull requests are displayed in this section, while any remaining pull requests beyond this limit are omitted for brevity.

Pull Requests Closed This Week: 197

Key Closed Pull Requests

1. Feat/qlora training: This pull request introduces a QLoRA fine-tuning example for Mixture-of-Experts (MoE) models, implementing the MUL_MAT_ID backward pass and adding support for both standard and reward-weighted supervised fine-tuning with gradient checkpointing to optimize VRAM usage, along with new diagnostic tools, sample datasets, and necessary backend operations and optimizer enhancements.

  • URL: pull/22705
  • Associated Commits: bcd40, d29de, 12df7, ee6d4, cb95b, 9b38c, 31edb, 6af45, 6c101, ebc77, cd1cb, fab18, 3760d, 8a30a, b1da2, 71350, 5ad16, e7a9c, b91ff, e28d7, e09e9, b3576, d956d, f70db, c0feb, 15dcf, b6e5a, ec3f5, 3f610, 8639f, c7404, baf09, 5b899, 9b29b, 07e85, 07656, e99da, 7f914, 9dd23, 452f8, 38692, 6bdad, adab8, 71d4f, f7338, 34304, 1ae4f, 800f4, 3f907, 79d8a, 65d87, 9af68, 24a10, 7bd29, 3bdb7, 40357, e4a51, 8bf60, 63857, 886e2, e73ad, 76e91, 3d76e, d9255, 02e75, d1c3d, 45378, e2ea5, 4d96f, e8001, 4d24b, 0c928, d8c66, 38211, 18e15, 9ed86, 75662, 05461, dd2a8, eb53b, ae567, f6d1c, d83d3, f1724, 056e6, 186c4, c7787, fe9ef, 3809c, 13b22, 170f8, 0cd2d, 6dca1, 6064e, 1e6bb, 48e85, 7f9af, 5fe35, d6cbf, fc940, e3270, e6e8f, af9b4, 6c984, 8416e, 0e6df, bfc51, f76ee, 7b732, f9687, 4f694, bd799, aabc2, 4d172, 8c34d, 1d026, c61a1, 512b9, 41d28, b503d, 5de24, 2988e, fae2c, 2b819, 6e21b, 2d4e4, 403ae, d535d, 88295, 15a2c, 80a75, f0459, f8038, aaba1, 4a5ba, 23da1, 2b11a, b1246, 39003, e16af, 64ae5, f0fea, 2d482, d1db6, 6616f, 0eb92, cf04d, 4737d, 28c09, 2d225, db887, a089b, e78be, 069b3, a9803, 639d1, 0296e, b2cca, 77cb3, ff2fd, f353e, e79f6, ad83f, f123c, 3ead4, 3560a, e01bf, 0eb09, 1549f, 72d60, d4836, 8772f, 40242, f2d14, 763e5, c1c16, 86540, fb401, 3063e, 47eff, f88b2, 7e7f9, 13c32, 7030d, 88ee0, 02299

2. hexagon: new HMX-optimized GATED_DELTA_NET: This pull request introduces a new HMX-optimized implementation of the GATED_DELTA_NET for the Hexagon backend, delivering prompt processing performance improvements of 1.5 to 3 times across various Qwen models and devices, and includes extensive enhancements to DMA support, buffer management, and pipeline optimizations to fully leverage 64-bit DMA addressing and HVX vectorization.

  • URL: pull/29199
  • Associated Commits: 9903c, 6b728, 60422, e48f5, 9a1fb, b8109, d4a31, f89cd, cac51, 6607a, 9b126, 542f8, 67bef, f2eb5, bf73e, 109e6, 3e963, 0c5cf, db1fa, f5a3c, f63af, 2a5d6, 7786d, 10952, 7f412, c1131, 136e4, e12a2, 57fc2, b6c7a, 55aa1, 6b957, 42bce, 3357f, d882d, ed854, 1740c, ffbf5, 26573, 49878, eafeb, 0aa2b, 55a05, dbe43, e5ee6, 18924, 25cdc, a0bac, 40c85, 7a37f, 7d116, 628f7, dc18a

3. vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4: This pull request implements an int8 quantized cooperative matrix (coopmat1) matrix multiplication shader optimized specifically for AMD RDNA3 and RDNA4 architectures, supporting multiple quantization formats and significantly improving performance in prompt processing benchmarks for models like Strix Halo and MoE, while including architecture-specific optimizations and disabling slower quant formats on RDNA4.

  • URL: pull/27952
  • Associated Commits: 13f6d, e086b, 5f034, 16256, 513f4, b121e, 2a878, 5553b, b453b, 14a60, 95d14, ad015, 31b9b, 23356, 9f069, 15b3c, 57b42, a9e0e, ef5c6, 2b9c5, 57dd3, ab6e8, 42a62, 18252, eb46b, d2a5a, ed294, fa37d, d8ed3, 3cd76, dc08f, ba840, 5037f, ab94f, 4695f, 22f58, 3985f, 09266, eeeac, 2ba1d, 9ef59, f2520, 0eab5, c30a8, 6421f, 04d7e, 584ae, e85ff, 8b194, 965e5

Other Closed Pull Requests

  • Model selector UI enhancements: This pull request adds advanced options to the model selector UI by introducing reusable submenus for sampling and penalties settings with inline editable inputs bound to the settings store. It improves user control by showing server defaults as placeholders and disables these submenus until a model is loaded.
    • pull/27243
  • Windows ARM64 native build support: This pull request enables native building on Windows ARM64 systems using MSVC's cl.exe compiler by modifying the CMake build flow to avoid requiring clang. It addresses MSVC-specific issues with NEON intrinsics and volatile keyword usage, improving compatibility and performance profiling while noting some intrinsic support gaps remain.
    • pull/28362
  • Batch processing API and implementation: This pull request introduces the llama_batch_ext feature, updating the API to support batch processing with tokens, embeddings, and state. It supersedes a previous staged PR and includes various implementation refinements and compatibility fixes.
    • pull/24669
  • DSpark speculative decoding support: This pull request adds support for DSpark speculative decoding with Gemma4 backbone drafts by integrating a new "dspark" architecture that extends DFlash speculative decoding with a low-rank Markov head and confidence-based draft pruning. It also provides model conversion, building, and running instructions.
    • pull/25549
  • W4A8 computation path on Blackwell GPUs: This pull request forces the W4A8 computation path for NVFP4_W4A16 layers on Blackwell GPUs by adding GGUF metadata, updating model hyperparameters, and applying a new dispatch hint including for MoE layers. It demonstrates improved model quality and stricter test thresholds as a result.
    • pull/24364
  • Hexagon backend quantization and kernel improvements: This pull request improves dynamic activation quantization by enhancing Q8_0 accuracy, supporting 64-bit DMA addresses, and replacing legacy DDR paths with new VTCM/DMA-based kernels to fix tolerance failures during LFM2.5-DSpark bring-up.
    • pull/29395
  • Hexagon backend sampling and chunking support: This pull request adds and updates support for backend sampling operations like ARGMAX, ARGSORT, TOP_K, SUM, and STEP, introduces chunking for large logit tensors, updates supported operations including ABS and RELU, and includes multiple fixes to improve feature parity and power efficiency.
    • pull/29502
  • Vulkan cooperative matrix support on Qualcomm Adreno GPUs: This pull request enables the VK_KHR_cooperative_matrix Vulkan extension on Qualcomm Adreno GPUs with hardware matrix cores by adding architecture detection, conditional activation, and Adreno-specific medium warptile configurations to optimize performance.
    • pull/29328
  • Per-sequence state restore cleanup fixes: This pull request ensures incomplete cleanup after failed per-sequence state restores by clearing affected tensor data and discarding deferred writes after rollback. It adds regression tests to verify that failed restores remove corrupted state without affecting successful operations.
    • pull/27530
  • MUSA backend fixes and improvements for PH1 platform: This pull request implements multiple fixes including enabling GATED_DELTA_NET, correcting quantized matrix multiplication, fixing dequantization kernels, improving flash attention, updating device initialization, enabling CUB library paths, and adjusting CI/build configurations for full PH1 support with verified correctness and performance gains.
    • pull/29193
  • Graph input reporting and debugging enhancements: This pull request enhances llama-context.cpp by adding detailed reporting of graph inputs and input tensors during sched_reserve, correcting batch-size labels, logging warnings for unexpected input tensor operations, and providing trace logs listing nodes using each input tensor to improve debugging.
    • pull/26625
  • MoE gradient computation and optimizer extensions: This pull request introduces gradient computation support for the MUL_MAT_ID operation in MoE models, extends the AdamW optimizer with per-element gradient clipping, and fixes backward graph builder issues to enable fine-tuning of MoE architectures.
    • pull/22704
  • Multi-host binding for llama-server: This pull request adds support for binding llama-server to multiple IP addresses and Unix socket paths simultaneously by accepting a comma-separated list with the --host option. It manages multiple listeners sharing the same API and model via a shared HTTP worker pool, allowing access through VPN and localhost.
    • pull/28690
  • REAP-style expert profiling and pruning tools: This pull request introduces C++ profiling tools that collect saliency scores from quantized GGUF models without full VRAM requirements and Python pruners that reduce model size by slicing expert weight tensors while preserving quantization, enabling significant compression with minimal quality loss.
    • pull/20454
  • Intel Xe Vulkan flash attention optimizations: This pull request adds Vulkan flash attention kernels optimized for Intel Xe platforms (Xe-LPG Plus, Xe2, Xe3) with new shaders and pipeline configurations. It enables efficient cooperative matrix operations and subgroup-based softmax reductions, resulting in significant speedups demonstrated by benchmarking.
    • pull/24406
  • Metal FWHT kernel enhancements: This pull request introduces new Metal FWHT kernels supporting block widths from 1024 to 8192 by running one row per threadgroup instead of per simdgroup. It enables efficient handling of wider blocks with F32 and F16 source types while respecting device memory limits and maintaining compatibility with existing kernels.
    • pull/29095
  • Hexagon backend tiled GET_ROWS support: This pull request adds support for tiled (repacked) layout paths of Q4_0 and Q8_0 GET_ROWS operations on Hexagon backend, replacing CPU fallback with a tiled implementation. It improves performance on Snapdragon X2 Elite HTP0 devices for large vocabulary tensors in the LFM2.5-DSpark draft model.
    • pull/29511
  • Vulkan backend matrix multiplication fix: This pull request fixes incorrect results in Vulkan backend by correctly reading batch stride from tensor metadata during in-place matrix multiplications on strided views. It prevents infinite loops in the SheetSage2 decoder and aligns Vulkan behavior with CPU and CUDA backends.
    • pull/28956
  • Server CI workflow update: This pull request updates the Server (sanitize) CI workflow by switching the runner from hf-jobs-cpu-xl to hf-jobs-cpu-performance and configures server tests to run with 4 pytest workers to improve testing efficiency.
    • pull/29369
  • Test refactoring and stability improvements: This pull request refactors test-recurrent-state-rollback.cpp to use llama_context_ptr for automatic context management, adds a --models DIR mode for running tests across dummy models, separates test_multi_seq_split_replay into an independent test, fixes graph reallocation issues in Metal fusion, and loosens split replay NMSE bounds to reduce flakiness, all verified by extensive testing.
    • pull/29426
  • Intel Xe GEMM and shader optimizations: This pull request introduces GEMM and Group GEMM kernel optimizations, load-time weight compression, and shader improvements targeting Intel Xe platforms to enhance performance on the Intel quick MoE path. It includes SLM-based matrix layout changes, dequantization optimizations, alternative pipeline selection, and fused kernel dispatches, resulting in significant speedups on Intel Arc GPUs.
    • pull/24407
  • Fused QKV tensor splitting fix: This pull request fixes tensor splitting logic for fused query-key-value operations to correctly handle uneven key and value head sizes, improving model compatibility and performance.
    • pull/29294
  • Backend-specific testing option: This pull request adds a -b/--backend option to the test-llama-archs tool, allowing users to run tests for a specific backend to facilitate targeted and efficient local testing.
    • pull/27372
  • Muse Glimmer parser error fix: This pull request fixes a parser error in Muse Glimmer that occurred when a turn started immediately with a tool call without a preceding reasoning step, aligning behavior with GPT-OSS and including a test to validate the fix.
    • pull/29242
  • ggml library synchronization: This pull request updates the project to synchronize with ggml version 0.25.0, including fixes for SYCL compile warnings and updates to release summary tasks.
    • pull/29300

3.3 Pull Request Discussion Insights

This section will analyze the tone and sentiment of discussions within this project's open and closed pull requests that occurred within the past week. It aims to identify potentially heated exchanges and to maintain a constructive project environment.

Based on our analysis, there are no instances of toxic discussions in the project's open or closed pull requests from the past week.


IV. Contributors

4.1 Contributors

Active Contributors:

We consider an active contributor in this project to be any contributor who has made at least 1 commit, opened at least 1 issue, created at least 1 pull request, or made more than 2 comments in the last month.

If there are more than 10 active contributors, the list is truncated to the top 10 based on contribution metrics for better clarity.

Contributor Commits Pull Requests Issues Comments
ggerganov 218 22 0 43
max-krasnyansky 169 4 0 4
allozaur 120 7 0 0
CISC 65 13 0 40
ServeurpersoCom 87 11 0 13
105 0 0 0
ngxson 71 6 0 23
aldehir 80 1 0 4
edwinbrowwn 84 0 0 0
wanghqc 71 0 0 0

Don't miss what's next. Subscribe to Weekly Project News:
Powered by Buttondown, the easiest way to start and grow your newsletter.