Weekly GitHub Report for Xla: September 21, 2026 - September 28, 2026 (20:38:15)
Weekly GitHub Report for Xla
Thank you for subscribing to our weekly newsletter! Each week, we deliver a comprehensive summary of your GitHub project's latest activity right to your inbox, including an overview of your project's issues, pull requests, contributors, and commit activity.
Table of Contents
I. News
1.1 Recent Version Releases:
No recent version releases were found.
1.2 Version Information:
Please provide the version release information you would like me to analyze and summarize.
II. Issues
2.1 Top 5 Active Issues:
We consider active issues to be issues that that have been commented on most frequently within the last week. Bot comments are omitted.
-
[BUG] [STAT:AWAITING OPENXLA-ENG] [NVIDIA-GPU] [GPU] [ERR: RUNTIME] CUDA initialization fails on NVIDIA fractional vGPU profiles without VMM support: This issue reports a regression in CUDA initialization for NVIDIA fractional vGPU profiles that lack Virtual Memory Management (VMM) support, causing XLA to fail during device initialization with an internal error. The user requests that XLA either restore fallback support for non-VMM devices or provide a clear opt-out mechanism and error messaging, as the current behavior breaks compatibility with hardware configurations that previously worked in earlier jaxlib versions.
- The comments acknowledge the need for improved error messaging and clarify that MIG partitioning differs from fractional vGPUs, with VMM available on MIG but not on fractional vGPUs. They also seek confirmation on whether non-VMM vGPU profiles are intentionally unsupported and discuss potential solutions, including pinning to older jaxlib versions or restoring fallback allocator support behind a flag.
- Number of comments this week: 2
-
[STAT:AWAITING OPENXLA-ENG] [ERR:PERFORMANCE] ConvertElementType bool-to-int returns incorrect result for nonstandard booleans: This issue addresses a problem in the StableHLO specification implementation where converting nonstandard boolean values to integers using ConvertElementType returns incorrect results, specifically failing to convert true values consistently to one. The discussion highlights challenges in fixing this behavior across different backends, the difficulty of supporting bitcast conversions for boolean types in XLA, and the implications for JAX's numpy view functionality, with recent progress made on CPU and GPU but awaiting TPU support.
- The comments detail attempts to fix the conversion logic, note compiler optimization issues, discuss related numpy behavior, and explore extending XLA's bitcast conversion to support booleans; progress has been made on CPU/GPU backends, with TPU support pending, and a recent commit removes blockers for JAX to adjust its workaround.
- Number of comments this week: 1
-
[BUG] [STAT:AWAITING OPENXLA-ENG] [ERR: RUNTIME] XLA tf.math.maximum has inconsistent behaviour in eager and compile modes when comparing NaN: This issue reports that the TensorFlow function
tf.math.maximumexhibits inconsistent behavior when comparingNaNvalues between eager execution and XLA compilation modes, specifically with tensors of float dtype. The user provides a minimal reproducible example demonstrating that the output differs in the presence ofNaNvalues, which could lead to unexpected results in computations involvingNaN.- An internal reproducer for the issue has been submitted, indicating that the problem is acknowledged and under investigation by the development team.
- Number of comments this week: 1
-
[BUG] [STAT:AWAITING OPENXLA-ENG] [ERR: RUNTIME] tf.math.igammac produces inconsistent results between eager execution and XLA compilation for negative inputs: This issue reports that the TensorFlow function
tf.math.igammacproduces inconsistent results when given negative inputs, yielding NaN values during eager execution but returning 1.0 during XLA compilation. This discrepancy suggests a potential bug in how negative inputs are handled differently between the eager and XLA execution paths.- An internal reproducer has been submitted to help investigate and address the inconsistency between eager execution and XLA compilation results for negative inputs in
tf.math.igammac. - Number of comments this week: 1
- An internal reproducer has been submitted to help investigate and address the inconsistency between eager execution and XLA compilation results for negative inputs in
-
[BUG] [STAT:AWAITING OPENXLA-ENG] [CPU] [ERR: RUNTIME] [XLA:CPU] In-place dynamic-update-slice fusion reads its own output when the update is a shifted slice of the operand: This issue describes a bug in the CPU backend of the XLA compiler where an in-place dynamic-update-slice operation fused with a shifted slice of the operand produces incorrect results due to the fusion reading from its own output buffer. The problem arises because the CPU fusion emitter incorrectly allows the fusion to proceed in-place, causing the kernel to read overwritten elements, whereas GPU execution and disabling certain optimization passes produce the correct output.
- The single comment notes that an internal reproducer for the issue has been submitted, indicating that the problem has been acknowledged and is being investigated further.
- Number of comments this week: 1
2.2 Top 5 Stale Issues:
We consider stale issues to be issues that has had no activity within the last 30 days. The team should work together to get these issues resolved and closed as soon as possible.
As of our latest update, there are no stale issues for the project this week.
2.3 Open Issues
This section lists, groups, and then summarizes issues that were created within the last week in the repository.
Issues Opened This Week: 31
Summarized Issues:
- XLA compilation failures due to shape and dtype mismatches: Multiple issues describe failures during XLA compilation caused by shape mismatches, dtype mismatches, or invalid broadcasting. These problems often occur despite successful eager execution, indicating discrepancies in how XLA handles dynamic shapes and type propagation compared to eager mode.
- issues/49185, issues/49196, issues/49332, issues/49334, issues/49337, issues/49516, issues/49517, issues/49529, issues/49617
- Runtime crashes and assertion failures in XLA and LLVM components: Several issues report runtime assertion failures, segmentation faults, or crashes triggered by invalid index accesses or incorrect assumptions in LLVM or XLA internals. These failures cause program aborts or core dumps during execution or compilation.
- issues/49222, issues/49223, issues/49227, issues/49544
- Discrepancies between eager execution and XLA behavior: Multiple issues highlight behavioral differences where operations succeed in eager mode but fail or behave incorrectly under XLA compilation or execution. These include errors like UnimplementedError, InternalError, or silent shape divergences.
- issues/49242, issues/49248, issues/49328, issues/49519, issues/49534, issues/49539, issues/49562, issues/49563, issues/49564
- Bugs in XLA SPMD partitioner and backend implementations: Some issues describe bugs in the SPMD partitioner causing incorrect all-reduces or crashes due to sharding mismatches, as well as backend-specific bugs like incorrect calling conventions or custom call failures. These bugs lead to incorrect results or crashes during partitioning or code generation.
- issues/49382, issues/49386, issues/49479, issues/49515
- Memory leaks and resource exhaustion during compilation: One issue reports unbounded host memory growth leading to out-of-memory termination when compiling with the fusion autotuner enabled on GPU, which is resolved by disabling the autotuner. This indicates resource management problems during compilation.
- issues/49549
- Failures related to dynamic shapes and bounded dynamic dimensions: Several issues involve failures or incorrect behavior when handling tensors with dynamic or bounded dynamic shapes, especially in operations like tf.unique, tf.squeeze, tf.repeat, and broadcasting. These problems cause compilation errors or incorrect outputs under XLA.
- issues/49185, issues/49516, issues/49517, issues/49562
- Errors in tensor operations involving complex64 and float16 types: Some issues describe incorrect type handling or invalid results when using complex64 or float16 tensors, including incorrect lowering, infinite results, or dtype mismatches that cause compilation failures or runtime errors.
- issues/49332, issues/49534, issues/49539
- Failures in XLA compilation due to control flow and nested loops: One issue describes a compilation failure caused by the DynamicSliceAnnotator pass when a dynamic slice offset depends on a nested while loop inside a lax.scan operation, indicating difficulties in analyzing dependencies in nested control flow.
- issues/49299
- CUDA and GPU initialization and compatibility issues: One issue reports that CUDA initialization fails on NVIDIA fractional vGPU profiles lacking VMM support, causing XLA to reject these devices and preventing JAX from running on such hardware, which worked previously.
- issues/49252
- Errors in tensor operations involving broadcasting and shape propagation: Some issues describe incorrect broadcasted output shapes or loss of static shape information during XLA compilation, leading to errors in subsequent operations like tf.squeeze or subtraction.
- issues/49337, issues/49529
- Incorrect handling of NaN and special values in XLA: One issue highlights that NaN inputs are preserved during eager execution but replaced with zeros during XLA compilation in a quantization gradient function, causing output discrepancies.
- issues/49613
2.4 Closed Issues
This section lists, groups, and then summarizes issues that were closed within the last week in the repository. This section also links the associated pull requests if applicable.
Issues Closed This Week: 10
Summarized Issues:
- Symbol collisions and PCI device enumeration issues: XLA's bundled hwloc library exports global symbols without PCI discovery support, which causes symbol collisions that break third-party libraries like aws-ofi-nccl and NIXL on AWS EFA instances. This prevents proper PCI device enumeration, leading to errors and performance degradation.
- issues/39355
- GPU kernel and floating-point comparison bugs: The XLA GPU TopK custom kernel for k ≤ 16 uses raw C++ floating-point comparisons instead of the IEEE 754 TOTALORDER, causing incorrect and corrupted output when -NaN values are present. Additionally, the GPU lowering of
dynamic-update-sliceoperations incorrectly computes byte offsets from unclamped start indices, resulting in stale values and incorrect data placement. - issues/48119, issues/49120
- Incorrect padding folding in AlgebraicSimplifier: Folding a negative (cropping) pad into the window padding of a consuming reduce_window or convolution operation merges padding values incorrectly. This causes jit-compiled programs to produce incorrect results by using input elements where padding values should be.
- issues/48715
- Non-deterministic GPU compilation and fusion priority: GPU compilation in XLA is non-reproducible due to non-deterministic iteration order over a hash set in GPU HLO cost analysis, leading to variations in fusion priority decisions. This results in different optimized HLO, PTX code, and floating-point results across compilations, fixed by enforcing deterministic iteration order sorted by unique IDs.
- issues/49047
- Compilation failures due to type mismatches in tensor operations: XLA compilation fails with InvalidArgumentError when applying tf.sign to complex64 outputs from tf.raw_ops.SelfAdjointEigV2, causing a type mismatch during DivNoNan between float32 and complex64 tensors. Similarly, using tf.experimental.numpy.min on zero-dimensional tf.keras.layers.Embedding outputs causes dimension misinterpretation during reduction, leading to compilation failure despite successful eager execution.
- issues/49128, issues/49169
- Miscompilation in gradient accumulation on TPU: The
while-loop-all-reduce-code-motionoptimization pass in JAX causes incorrect gradient accumulation during scatter operations within ascanon TPU devices. This results in significant numerical errors compared to NumPy references and correct CPU execution, resolved by disabling this optimization. - issues/49143
- Inconsistent GPU topology fingerprinting across processes: The GPU topology fingerprint generated by XLA is process-local and depends on the first local device's interconnect information, causing inconsistent fingerprints across processes in multi-GPU jobs. This leads to problems with JAX's persistent compilation cache and cross-process program compilation synchronization.
- issues/49371
- ScatterDeterminismExpander bug with out-of-bounds indices: Under deterministic GPU operations, the ScatterDeterminismExpander incorrectly writes out-of-bounds scatter indices with scalar updates into the next row instead of dropping them. This violates the StableHLO specification and causes unexpected behavior in JAX's scatter operation with mode "drop".
- issues/49380
2.5 Issue Discussion Insights
This section will analyze the tone and sentiment of discussions within this project's open and closed issues that occurred within the past week. It aims to identify potentially heated exchanges and to maintain a constructive project environment.
Based on our analysis, there are no instances of toxic discussions in the project's open or closed issues from the past week.
III. Pull Requests
3.1 Open Pull Requests
This section provides a summary of pull requests that were opened in the repository over the past week. The top three pull requests with the highest number of commits are highlighted as 'key' pull requests. Other pull requests are grouped based on similar characteristics for easier analysis. Up to 25 pull requests are displayed in this section, while any remaining pull requests beyond this limit are omitted for brevity.
Pull Requests Opened This Week: 14
Key Open Pull Requests
1. Test new ML Actions Runners: This pull request updates the continuous integration workflows and benchmark configurations to test new machine learning Actions Runners, including adding pull request triggers to nightly benchmarks, introducing new Linux CPU configurations, refactoring CI jobs, and enabling manual and pull request triggers for benchmarks to improve testing coverage and performance validation.
- URL: pull/49443
- Associated Commits: 65b53, d7fff, 96571, a6583, 8fcb9, ab23d, f020b, 09d45, 5a24c, 18e14, c7437, 2489a, e195e, 6450a
2. [XLA:CPU] Add CpuPerformanceModel: This pull request introduces a CpuPerformanceModel and CpuHloCostAnalysis to provide a CPU-specific cost analysis and performance modeling framework that improves fusion decision heuristics by accounting for memory caching behavior and broadcast operand handling, addressing issues with current CPU fusion strategies that misclassify memory-bound instructions as expensive.
- URL: pull/49546
3. [ROCm] Move shim script logic upstream: This pull request moves the ROCm shim script logic upstream to streamline the release branch creation process by minimizing pipeline changes, ensuring the same pipeline functions consistently across ROCm release branches.
- URL: pull/49495
Other Open Pull Requests
- Bug fixes in XLA operations and caching: Several pull requests address critical bugs in XLA components, including fixing batch dimension handling in the
DotDecomposerto prevent compilation failures, resolving ROCm matmul plan cache thrashing by relaxing caching requirements, and improving GPU kernel compilation by inlining scan computation calls to fix indexing map errors. These fixes enhance stability and correctness across CPU and GPU backends with added unit tests to verify the changes. - pull/49346, pull/49300, pull/49603
- Enhancements to GPU backend and autotuning: Improvements include supporting table offsets for dynamic slices in the GPU backend to handle non-linear sequences while maintaining backward compatibility, introducing autotuning of MFMA instruction sizes in Triton GEMM configurations for better performance on MI300 and MI350 GPUs, and optimizing the GPU autotuner compare process by moving mismatch counters to device memory and parallelizing host-side comparisons. These changes significantly boost performance and efficiency in GPU computations.
- pull/49436, pull/49450, pull/49601
- Timeout and watchdog improvements for XLA:GPU: A new approach separates host and device execution timeouts by replacing the
ExecutionWatchdogScopewith a scopedHangWatchdogfor host monitoring and adding a device execution timeout parameter that is disabled by default. This update clarifies timeout semantics and better handles CUDA API deadlocks and kernel hangs. - pull/49276
- Runtime custom options and cross-host GPU transfer improvements: One pull request exposes per-execution runtime custom options to FFI handlers, enabling independent configuration of executions beyond static attributes, while another enhances cross-host GPU transfer by incorporating task incarnation information into clique keys to ensure distinct keys for restarted tasks and improve transfer consistency. Both changes improve configurability and robustness in distributed GPU environments.
- pull/49360, pull/49401
- Eigen library update for numerical stability: The Eigen dependency is updated to include upstream fixes that address underflow and overflow issues in the
makeHouseholderfunction, improving numerical stability in TensorFlow's CPUtf.linalg.qrkernel by preventing loss of significance and NaN results in float32 vectors. - pull/49333
3.2 Closed Pull Requests
This section provides a summary of pull requests that were closed in the repository over the past week. The top three pull requests with the highest number of commits are highlighted as 'key' pull requests. Other pull requests are grouped based on similar characteristics for easier analysis. Up to 25 pull requests are displayed in this section, while any remaining pull requests beyond this limit are omitted for brevity.
Pull Requests Closed This Week: 40
Key Closed Pull Requests
1. [oneDNN][CPU] Add initial FP8 support for oneDNN dispatch: This pull request introduces initial support for FP8 data types in oneDNN CPU dispatch by adding detection and gating of the AMX-FP8 CPU feature with an AMX-FP16 emulation fallback, extending CPU feature detection, updating dispatch logic, and adding corresponding tests to enable calling oneDNN for FP8 functionality on x86 CPUs.
- URL: pull/43318
2. Test branch v10: This pull request is about a series of ROCm-related fixes and improvements including streamlining bazel targets by moving from dynamic loading to direct linking of ROCm libraries, unblocking CI builds, adding support for packed bf16 atomic add operations to improve performance on MI300/MI350 GPUs, fixing compatibility issues with GCC 13 in device code, updating device query tools from rocm-smi to amd-smi, gating collective kernels based on hardware atomic visibility to prevent hangs on certain GPUs, and various cleanup and test adjustments to enhance stability and performance of the ROCm backend in the XLA project.
- URL: pull/49482
3. [DO NOT MERGE] CI check: This pull request is about adding and refining continuous integration (CI) checks for the ROCm platform, including setting the ROCM_PATH environment variable to fix build issues, improving GPU sandboxing on GitHub Actions runners, managing Docker configurations and logins, and testing dynamic mode settings, although it was ultimately not merged.
- URL: pull/47243
Other Closed Pull Requests
- GPU memory optimization and fusion improvements: Multiple pull requests focus on optimizing GPU memory usage by rematerializing copies of large constant fill operations as independent clones or rewriting copies of large constant-fill fusion operations into independent fills. These changes enable separate scheduling, shorter buffer lifetimes, and reduced peak memory usage without affecting correctness or performance.
- pull/49306, pull/49376
- Fusion cost modeling and CPU fusion enhancements: A pull request introduces a CPU fusion cost model and related improvements to enhance fusion decisions in the XLA compiler. This enables more efficient fusion of producers into consumers by considering machine-specific cost metrics, reducing unnecessary intermediate buffers, and improving performance on reduction and pairwise workloads.
- pull/49316
- Bug fixes in GPU backend and execution ordering: Several pull requests address bugs in the XLA GPU backend, including preserving asynchronous operation ordering across command buffer boundaries and rewriting FP8 convolutions to optimized calls in all computations. These fixes improve correctness, prevent incorrect reordering, and enhance performance on FP8-capable GPUs.
- pull/49149, pull/49007
- Bug fixes and safety improvements in XLA passes and simplifiers: Pull requests fix bugs in the AlgebraicSimplifier related to negative padding folding and enhance the while-loop all-reduce code motion pass to correctly handle predicate replication in scatter carry patterns. These changes prevent silent incorrect results and ensure unsafe scatter patterns retain correct all-reduce placement.
- pull/49351, pull/49182
- Custom execution options forwarding and runtime enhancements: One pull request introduces a mechanism to forward custom execution options from the IFRT layer through the PjRt interface to XLA CPU and GPU runtimes. This enables per-call runtime-defined attributes to be passed without retracing executables by extending execution option structures and updating APIs accordingly.
- pull/48829
- Memory allocator reconfiguration for GPU: A pull request reconfigures the shared BFC memory allocator in XLA to anchor collective memory allocations at the lower end of the shared arena and default memory at the upper end. This preserves distinct splitting and hole reuse policies for each memory space, ensuring stable offsets for collective buffers and preventing capacity regressions during arena growth.
- pull/48834
- Dependency and backend updates: Pull requests add the oneDPL 2022.13.0 dependency as a hermetic Bazel third-party library with stub implementations for SYCL oneAPI backend, introduce the mi450 backend replacing gfx1250, and update the project to use the newest rules_ml_toolchain with hermetic ROCm download overrides.
- pull/47640, pull/48355, pull/49152
- Bug fixes and improvements in CPU backend and code safety: Pull requests replace undefined behavior caused by strict-aliasing violations with
absl::bit_cast, add anempty()predicate to simplify zero-size slice checks, and fix MSVC math constants inclusion by adding<cmath>after defining_USE_MATH_DEFINES. These changes improve code safety, readability, and correctness. - pull/49146, pull/49148, pull/49149
- Bug fixes in precision handling and operand swapping: A pull request updates GemmFusionSwapOperands to correctly swap operand precision entries along with operands, ensuring precision settings align with swapped inputs and adding tests for different precision configurations.
- pull/49110
- Bug fix for XLA GPU execution HangWatchdog: One pull request adds a second timeout watch that aborts the process if the first timeout handler fails to unwind, preventing indefinite hangs and enabling system schedulers to restart jobs.
- pull/49168
- Symbol renaming to fix PCI count issues: A pull request renames bundled hwloc symbols with a
tf_prefix in third-party builds to prevent symbol interposition with system PCI discovery, fixing PCI count issues on AWS EFA and ensuring proper linkage of split DSOs. - pull/47975
- Explicit disabling of oneDNN custom calls for tests: A pull request disables oneDNN custom calls for YNN-specific tests to prevent unit test failures when oneDNN is disabled by default.
- pull/48307
- Explicit error handling for unsupported ROCm collective features: A pull request adds explicit
UnimplementedErrorin collective info builders when cross-host one-shot collective is enabled on ROCm targets and marks related tests as CUDA-only due to lack of symmetric memory support on ROCm. - pull/48586
- Preliminary build and CI updates: Pull requests include a preliminary build check with minor fixes not ready for merging and updates to README, test runner, and continuous integration workflows.
- pull/49062, pull/23365
3.3 Pull Request Discussion Insights
This section will analyze the tone and sentiment of discussions within this project's open and closed pull requests that occurred within the past week. It aims to identify potentially heated exchanges and to maintain a constructive project environment.
Based on our analysis, there are no instances of toxic discussions in the project's open or closed pull requests from the past week.
IV. Contributors
4.1 Contributors
Active Contributors:
We consider an active contributor in this project to be any contributor who has made at least 1 commit, opened at least 1 issue, created at least 1 pull request, or made more than 2 comments in the last month.
If there are more than 10 active contributors, the list is truncated to the top 10 based on contribution metrics for better clarity.
| Contributor | Commits | Pull Requests | Issues | Comments |
|---|---|---|---|---|
| alekstheod | 66 | 4 | 0 | 0 |
| shawnwang18 | 38 | 7 | 0 | 2 |
| ezhulenev | 13 | 4 | 0 | 15 |
| jasminetrail | 0 | 0 | 20 | 0 |
| jurasic-pf | 14 | 4 | 1 | 0 |
| akuegel | 0 | 0 | 0 | 17 |
| quoctruong | 14 | 1 | 0 | 0 |
| akhilgoe | 8 | 4 | 0 | 0 |
| kodlan | 10 | 1 | 0 | 0 |
| ScXfjiang | 11 | 0 | 0 | 0 |