Weekly GitHub Report for Keras: August 31, 2026 - September 07, 2026 (21:21:22)
Weekly GitHub Report for Keras
Thank you for subscribing to our weekly newsletter! Each week, we deliver a comprehensive summary of your GitHub project's latest activity right to your inbox, including an overview of your project's issues, pull requests, contributors, and commit activity.
Table of Contents
I. News
1.1 Recent Version Releases:
The current version of this repository is v3.14.0
1.2 Version Information:
Released on April 2, 2026, this version introduces full Orbax checkpoint integration with advanced quantization methods like AWQ and INT4 sub-channel quantization, a new ScheduleFreeAdamW optimizer, and optional Gated Attention in key attention layers. It also significantly enhances OpenVINO backend support with extensive NumPy, neural network, and control flow operations, adds numerous new math and preprocessing features, and includes various backend-specific improvements and bug fixes.
Click here to view the full release notes!
II. Issues
2.1 Top 5 Active Issues:
We consider active issues to be issues that that have been commented on most frequently within the last week. Bot comments are omitted.
-
[TYPE:BUG] [BACKEND:TENSORFLOW] model.export() / ExportArchive store every variable twice in the SavedModel — 2× artifact size and 2× live variable RAM in every consumer (incl. TF Serving): This issue describes a problem in the TensorFlow backend where exporting a model using
model.export()or the documentedExportArchivepattern results in every model weight being stored twice in the SavedModel checkpoint, causing approximately double the artifact size on disk and double the live variable memory usage in any process that loads the model, including TensorFlow Serving. The root cause is that the export process tracks both the Keras variable wrappers and the raw TensorFlow variables separately, which prevents TensorFlow's checkpoint de-duplication from consolidating them, leading to duplicated storage and live resource buffers.- The comments confirm the issue and discuss attempted fixes, including deduplicating variables by their underlying handles, but find that removing endpoint-captured variables breaks LiteRT exports. Instead, the consensus is to preserve endpoint variables and deduplicate by replacing the default tracked Keras variables with the endpoint-captured TensorFlow variables, which resolves duplication without breaking existing functionality; this approach has been tested and accepted by the contributors.
- Number of comments this week: 6
-
[TYPE:BUG] [BACKEND:TORCH] Bug: keras.src.ops.normalize() produces -inf gradients for small non-zero float16 inputs for the torch backend: This issue reports a bug in the
keras.src.ops.normalize()function where, for small non-zerofloat16inputs using the torch backend, the L2 normalization path produces-infgradients due to underflow and overflow in thefloat16dtype during squared norm and reciprocal square root computations. The user suggests a fix involving upcasting inputs tofloat32for these calculations to avoid non-finite gradients, ensuring stable and finite gradient values similar to those observed withfloat32inputs.- The comment section confirms the bug is real and not specific to the torch backend, discusses related past fixes that did not address this half-precision issue, and agrees on the proposed solution to upcast to
float32during computation, with plans to open a PR including a test forfloat16gradients. - Number of comments this week: 2
- The comment section confirms the bug is real and not specific to the torch backend, discusses related past fixes that did not address this half-precision issue, and agrees on the proposed solution to upcast to
-
[TYPE:FEATURE] [KERAS-TEAM-REVIEW-PENDING] 🚀 Contributing to Keras: Open Projects & Ideas 🚀: This issue serves as a comprehensive invitation for contributors to engage with the Keras 3 project by tackling various open tasks such as performance optimization, backend development, adding pretrained models, creating new tutorials, and increasing test coverage. It outlines specific areas where help is needed and provides guidance on how to get started, encouraging community involvement to enhance the Keras ecosystem.
- The comments reflect active community engagement with contributors volunteering for specific tasks, seeking clarifications on contribution processes, discussing implementation details, and receiving guidance on best practices and next steps from maintainers.
- Number of comments this week: 1
-
[TYPE:FEATURE] [Contribution Proposal] Adding a PaddlePaddle backend for Keras 3 – seeking community feedback: This issue is a proposal to add a PaddlePaddle backend to Keras 3, aiming to enable Keras users to run their code on PaddlePaddle in addition to existing backends like TensorFlow, JAX, and PyTorch. The contributor offers to develop and maintain this backend independently, seeking community feedback on technical requirements, integration approach, and potential concerns before starting the implementation.
- Multiple community members expressed interest in contributing or supporting the effort, discussed a phased implementation approach starting with basic BERT model support, and suggested creating a dedicated PaddlePaddle branch; the contributor shared contact information for further discussion and has already prepared an initial backend structure repository.
- Number of comments this week: 1
-
[TYPE:BUG] torch conv does two NHWC<->NCHW copies when one would do: This issue addresses an inefficiency in the torch backend's convolution operation where two redundant data copies occur during the forward pass when using the default
channels_lastdata format. The proposed fix involves merging the two separate tensor contiguity operations into a single step to eliminate unnecessary copying, thereby optimizing performance for 4-D and 5-D convolution paths.- The single comment indicates a user has taken ownership of the issue, suggesting they will work on or track the fix.
- Number of comments this week: 1
2.2 Top 5 Stale Issues:
We consider stale issues to be issues that has had no activity within the last 30 days. The team should work together to get these issues resolved and closed as soon as possible.
As of our latest update, there are no stale issues for the project this week.
2.3 Open Issues
This section lists, groups, and then summarizes issues that were created within the last week in the repository.
Issues Opened This Week: 6
Summarized Issues:
- Numerical stability and dtype issues: Several issues highlight problems with numerical computations due to improper dtype handling or lack of safeguards. The
keras.ops.hard_sigmoidfunction produces incorrect results for integer inputs because constants are not cast properly, and thekeras.src.ops.normalize()function causes negative infinity gradients with float16 inputs due to underflow and overflow, suggesting the need for upcasting. Additionally, thecategorical_crossentropyloss function suffers from division-by-zero NaNs when predictions sum to zero, requiring an epsilon for stability. - issues/23544, issues/23546, issues/23551
- Model export and serialization inefficiency: The TensorFlow Keras model export process duplicates every model variable by serializing it both as a
keras.Variableand a rawtf.Variable. This duplication leads to doubled artifact size on disk and increased live memory usage when loading the SavedModel, caused by distinct trackable objects wrapping the same buffers and preventing checkpoint de-duplication. - issues/23553
- New feature proposal for attention layers: A backend-agnostic, serializable
keras.layers.LinearAttentionlayer is proposed to implement kernelized self-attention with linear sequence complexity and constant-size state for causal streaming. This layer would support causal and non-causal modes, 2D sequence masks, mixed precision, dynamic sequence lengths, and cross-backend execution across TensorFlow, JAX, and PyTorch. - issues/23555
- API export inconsistency: The
randomsubmodule is not exported in the public API ofkeras.opsin versions 3.14.x and 3.15.x, causing anAttributeErrorwhen users try to access random functions viakeras.ops.random.*. Although the implementation exists internally and works throughkeras.src.ops, the missing export breaks expected public API access. - issues/23575
2.4 Closed Issues
This section lists, groups, and then summarizes issues that were closed within the last week in the repository. This section also links the associated pull requests if applicable.
Issues Closed This Week: 8
Summarized Issues:
- Inconsistent dtype handling in keras.ops: Several operations in
keras.opsexhibit inconsistent dtype behavior, such as silently downcasting float64 inputs to float32 or returning incorrect numerical results like infinity for small float64 inputs. This inconsistency causes confusion and incorrect outputs, indicating a need to clarify and unify dtype promotion and numerical behavior across operations. - issues/23223, issues/23401
- Inefficiency in Layer call attribute setting: The
Layer.__call__method redundantly re-assignsself.builtandself._calledtoTrueon every call, causing unnecessary overhead especially in PyTorch backends due to repeated attribute setting. A fix is proposed to guard these assignments so they only occur when the values actually change, improving efficiency. - issues/23280
- Documentation improvement for image loading: There is a need to improve documentation to inform developers that
keras.utils.load_img()inherits Pillow's decompression bomb protection, which raises aPIL.Image.DecompressionBombErrorwhen loading images exceeding the pixel size limit. This clarification helps users understand potential exceptions during image loading. - issues/23317
- Adam optimizer advantages for MNIST training: The choice of the Adam optimizer for MNIST training is discussed, emphasizing its benefits such as adaptive learning rates, faster convergence, less manual tuning, and more reliable performance compared to SGD and RMSProp. This explanation clarifies why Adam is preferred in this context.
- issues/23370
- Sparsemax activation output correctness: The
keras.ops.sparsemaxandkeras.activations.sparsemaxfunctions produce outputs that do not sum to 1 across multiple backends due to an incorrect threshold calculation in the Euclidean projection onto the probability simplex. This bug causes incorrect activation values whenever the support contains more than one element, affecting the correctness of sparsemax outputs. - issues/23425
- Mixed precision inference support for mixed_float8: There is a need to enable support for mixed precision inference using the mixed_float8 data type in Keras, particularly to allow setting the dtype policy to "mixed_float8" for hardware compatibility like Nvidia L4 GPUs. Currently, an error occurs due to the lack of a
_float8_callmethod implementation in the Functional layer, blocking this functionality. - issues/23479
- Inconsistent integer input handling in keras.ops.gelu: The
keras.ops.gelufunction behaves inconsistently with integer inputs across backends: NumPy silently returns zeros due to improper dtype casting, JAX computes correct values, and Torch raises an error. This inconsistency highlights the need for uniform handling or error reporting for integer inputs in the NumPy backend. - issues/23528
2.5 Issue Discussion Insights
This section will analyze the tone and sentiment of discussions within this project's open and closed issues that occurred within the past week. It aims to identify potentially heated exchanges and to maintain a constructive project environment.
Based on our analysis, there are no instances of toxic discussions in the project's open or closed issues from the past week.
III. Pull Requests
3.1 Open Pull Requests
This section provides a summary of pull requests that were opened in the repository over the past week. The top three pull requests with the highest number of commits are highlighted as 'key' pull requests. Other pull requests are grouped based on similar characteristics for easier analysis. Up to 25 pull requests are displayed in this section, while any remaining pull requests beyond this limit are omitted for brevity.
Pull Requests Opened This Week: 14
Key Open Pull Requests
1. Implement gammaincc function in keras.ops: This pull request implements the keras.ops.gammaincc function, which computes the regularized upper incomplete gamma function element-wise for input tensors across multiple backends including NumPy, TensorFlow, PyTorch, and JAX, with the exception of OpenVINO.
- URL: pull/23561
2. Fix saved model variables: This pull request fixes the issue of duplicate variable storage during TensorFlow SavedModel export by ensuring that endpoint-captured tf.Variable objects are used for variable collections, preventing variables from being stored twice in the checkpoint while maintaining compatibility with LiteRT.
- URL: pull/23563
3. Fix -inf gradients in normalize() for float16 inputs (compute L2 norm in float32): This pull request fixes the issue of -inf gradients occurring in the normalize() function for float16 inputs by computing the L2 norm intermediates in float32 before casting back to the original dtype, thereby preventing underflow and overflow problems during gradient calculation and ensuring stable backward passes across all backends.
- URL: pull/23566
Other Open Pull Requests
- Test exclusion logic refinement: This pull request updates the test exclusion logic in the backend to match test node IDs exactly rather than by substring, ensuring that only the intended tests are skipped. This improves the reliability and clarity of test exclusions across different backends.
pull/23569
- Continuous integration workflow refactor: This pull request refactors the continuous integration setup by extracting the CPU test job from the main
actions.ymlworkflow into a standalone, reusablecpu_tests.ymlworkflow. This enables a single source of truth for CPU test logic with configurable timeouts and improved consistency across test runs.
pull/23572
- Integer input promotion in activations and math operations: These pull requests address issues where integer inputs to activation functions like
hard_sigmoidandhard_siluand operations like reciprocal and rsqrt were incorrectly cast or truncated. By promoting integer inputs to float types across multiple backends, they ensure consistent and correct computation outputs aligned with other frameworks.
pull/23545, pull/23574
- Fixes for numerical stability and precision in gradients and loss functions: These pull requests fix a division by zero NaN vulnerability in
categorical_crossentropyby adding epsilon to the denominator and correct L2 normalization gradient overflow by performing calculations in float32 before casting to float16. Additionally, a validation check was added in thekeras.losses.Huberconstructor to prevent degenerate loss values from non-positive delta parameters.
pull/23552, pull/23565, pull/23573
- Reduction operations shape preservation and backend consistency: This pull request fixes a bug where reductions with
axis=Noneandkeepdims=Trueincorrectly dropped input rank across multiple backends. It consolidates rank-restoration logic into a shared helper function, corrects OpenVINO and Torch implementations, and adds comprehensive tests to ensure consistent shape preservation for ten reduction operations.
pull/23571
- API consistency and export improvements: These pull requests update the Keras codebase by replacing all instances of
backend.is_tensorwithbackend.ops.is_tensorto align with the public API and export therandomsubmodule under the publickeras.opsAPI. They include updates to API generation, import guards, and unit tests to ensure consistent access to random operations via bothkeras.randomandkeras.ops.random.
pull/23578, pull/23580
- IAM permission check for Artifact Registry: This pull request adds a read-only IAM permission check for the ml-public-container Artifact Registry by implementing a numpy-backend test that calls
testIamPermissionsand reports granted permissions in pytest failures. This is intended to gather logs before closing the PR.
pull/23582
3.2 Closed Pull Requests
This section provides a summary of pull requests that were closed in the repository over the past week. The top three pull requests with the highest number of commits are highlighted as 'key' pull requests. Other pull requests are grouped based on similar characteristics for easier analysis. Up to 25 pull requests are displayed in this section, while any remaining pull requests beyond this limit are omitted for brevity.
Pull Requests Closed This Week: 56
Key Closed Pull Requests
1. [DO NOT REVIEW] integration testing branch for #22561 series (13 PRs combined): This pull request is an internal integration and benchmarking branch that combines 13 independent pull requests from the #22561 torch eager-overhead series to verify their clean composition without conflicts, ensure bit-exact numeric equivalence with the master branch, and run the full combined test suite, serving solely as a standing integration check without requesting review or merging.
- URL: pull/23257
- Associated Commits: c2b77, a9d06, fb8f4, f84a4, d02c5, 3b423, 2528f, b360a, 7b90b, 7e487, 9e124, fbda2, b0cc2, 674e8, 677fe, 8ae7d, a9245, a5c82, 19df2, a4713, 839a9, 0f1e3, 09f40, 4be28, c7594, 1b201, f1376, 4ad3a, 397ef, 2c203, fc0e7, 7240f, 18fc2, adcf5, c65fa, a5733, 785f2, 7a4a7, d97c4, 348b6, b0415, c1aab, 53fc1, 49e92, 59d47, 04be7, f19bc, 6b16e, 28bb8, 25c89, 9b7ab, 48ebe, 7b78c, 0d414, abe7e, 8ac5e, d3cbe, 1c29c, 56882, 9793e, 94de1, 543f2, 0b792, ad278, 5b273, 500f5, 4f376, 2b77e, 0adca, 7a9cf
2. [Fix] keras.ops inconsistently downcasts float64 to float32: This pull request fixes the issue where keras.ops operations inconsistently downcast float64 inputs to float32 by introducing a shared dtype helper that ensures consistent handling of float64 arrays and lists across multiple affected operations such as add, subtract, multiply, divide, power, mod, and fmod when using the PyTorch backend.
- URL: pull/23225
- Associated Commits: cbc89, f93ba, 0a76d, 31922, 8b2bd, 6f7f0, 78622, 06b69, 1c2fc, d0063, 52e47, 127c6, 0120c, ff0ff, a7d2f, 98be9, 871d2
3. Refactor build process and enhance documentation: This pull request proposes a refactor of the build process to enhance clarity and maintainability, adds detailed docstrings to functions, improves handling of legacy imports, and includes various formatting and script updates to improve overall code quality and documentation.
- URL: pull/23448
Other Closed Pull Requests
- Backend-specific and agnostic implementations of lgamma operation: These pull requests add backend-specific implementations of the
lgammaoperation for JAX, TensorFlow, PyTorch, and Numpy, while maintaining a backend-agnostic implementation for OpenVino. This ensures consistent functionality of thelgammaoperation across different backends in Keras.
- Performance optimizations in Layer.call method: Multiple pull requests improve the performance of the
Layer.__call__method by introducing a fast path to skip costly signature binding and adding guards to avoid redundant attribute reassignments. These changes reduce overhead and eliminate graph breaks in compiled workloads, resulting in measurable speedups without changing functional behavior.
- Sparsemax activation function fixes: Several pull requests fix the sparsemax activation by correcting the threshold computation to sum sorted logits over the support instead of cumulative sums, resolving dtype promotion issues in the NumPy backend, and ensuring output dtype matches input dtype across backends. These fixes guarantee correct projection onto the probability simplex and consistent behavior across frameworks.
- Fixes and improvements to Keras ops and utilities: Pull requests address various fixes including the floating-point overflow in
keras.ops.rsqrtfor subnormal inputs, makingoutput_sizeoptional inkeras.ops.image.reconstruct_patcheswith validation and backward compatibility, and correctingkeras.utils.normalize()to preserve the shape of 1D NumPy inputs. These changes improve numerical stability, usability, and correctness of Keras operations.
- Fixes to torch backend and multiprocessing: One pull request fixes torch backend reductions to correctly honor
keepdims=Truewhenaxis=Nonefor various reduction operations, ensuring output shape consistency with numpy and jax. Another pull request fixes a segmentation fault in PyTorch multiprocessing by switching totorch.multiprocessingand enforcing thespawnstart method to prevent GPU memory conflicts.
- Model Parallel support for PyTorch backend: This pull request integrates Keras 3 ModelParallel distributions with PyTorch's DTensor infrastructure, enabling sharding of model weights and activations across multiple devices. It updates data distribution and tensor operations to be DTensor-aware, enhancing multi-device training capabilities.
- Fixes to layer and variable reassignment behavior: A pull request fixes and enables reassignment of sublayers and variables after a layer has been built by temporarily unlocking the tracker during reassignment. This aligns Keras behavior with PyTorch and TensorFlow Modules and removes the need for private workaround mechanisms.
- Security hardening of KerasFileEditor against HDF5 shape bombs: This pull request adds a cumulative size check to reject maliciously crafted weight files that can cause memory exhaustion, hardening the editor against memory attacks similar to those mitigated in the loader. It includes tests to ensure rejection of both single large datasets and multiple smaller datasets exceeding safe limits before any data is read.
- EarlyStopping callback improvements: This pull request improves the EarlyStopping callback by validating that the patience parameter is non-negative, resetting the wait counter on any improvement of the monitored metric, and adding regression tests covering edge cases like negative patience and sub-baseline improvements. These changes enhance robustness and correctness of early stopping behavior.
- Fixes to ops.gelu and convert_to_tensor function: One pull request fixes the
ops.gelufunction in the NumPy backend to correctly handle integer inputs by casting them to float32, resolving incorrect zero outputs. Another pull request fixes an always-true device comparison inconvert_to_tensorby removing ineffective checks, eliminating dead code and improving performance by about 10% on that path.
- Addition of core image restoration loss functions: A pull request implements five core image restoration loss functions as standalone functions and Keras Loss classes with full support for reduction, dtype, serialization, and the compile API. Comprehensive tests verify correctness, serialization, and backend compatibility across NumPy, TensorFlow, and JAX.
- Introduction of opt-in JAXEpochIterator for training performance: This pull request introduces a threaded JAXEpochIterator to potentially improve training iteration performance, with benchmarks showing modest speedups on various hardware when enabled via an environment variable.
- Fixes to in_top_k function for list/tuple inputs: This pull request fixes an AttributeError in
in_top_kwhen handling list or tuple inputs with rank greater than 2 by replacing a direct shape check withnp.ndim. The fix is applied across TensorFlow, NumPy, and JAX backends with tests to prevent regressions.
- MeanSquaredError test configuration improvements: This pull request replaces an empty test configuration placeholder with detailed configuration assertions and implements a serialization round-trip test to ensure correctness and robustness of the MeanSquaredError configuration.
- Read-only CI runner identity probe extension for Artifact Registry and GCR: This pull request extends the existing read-only CI runner identity probe to include calls to Artifact Registry and Google Container Registry testIamPermissions for verifying IAM permissions as part of authorized OSS VRP research. It performs no uploads or secret dumps and is intended to be closed after capturing relevant job logs without merging.
3.3 Pull Request Discussion Insights
This section will analyze the tone and sentiment of discussions within this project's open and closed pull requests that occurred within the past week. It aims to identify potentially heated exchanges and to maintain a constructive project environment.
Based on our analysis, there are no instances of toxic discussions in the project's open or closed pull requests from the past week.
IV. Contributors
4.1 Contributors
Active Contributors:
We consider an active contributor in this project to be any contributor who has made at least 1 commit, opened at least 1 issue, created at least 1 pull request, or made more than 2 comments in the last month.
If there are more than 10 active contributors, the list is truncated to the top 10 based on contribution metrics for better clarity.
| Contributor | Commits | Pull Requests | Issues | Comments |
|---|---|---|---|---|
| pctablet505 | 84 | 0 | 0 | 0 |
| SamanehSaadat | 24 | 4 | 0 | 13 |
| hertschuh | 7 | 5 | 1 | 24 |
| buildwithsuhana | 23 | 4 | 0 | 0 |
| Neilblaze | 14 | 5 | 2 | 2 |
| 18 | 0 | 0 | 0 | |
| JyotinderSingh | 16 | 1 | 0 | 0 |
| M0nd0R | 12 | 5 | 0 | 0 |
| MarcosAsh | 12 | 4 | 0 | 0 |
| MaddipatlaChetan24 | 13 | 2 | 0 | 0 |
Access Last Week's Newsletter: