Weekly Project News

Archives
Subscribe

Weekly GitHub Report for Xla: August 04, 2026 - August 11, 2026 (00:34:56)

Weekly GitHub Report for Xla

Thank you for subscribing to our weekly newsletter! Each week, we deliver a comprehensive summary of your GitHub project's latest activity right to your inbox, including an overview of your project's issues, pull requests, contributors, and commit activity.


Table of Contents

  • I. News
    • 1.1. Recent Version Releases
    • 1.2. Other Noteworthy Updates
  • II. Issues
    • 2.1. Top 5 Active Issues
    • 2.2. Top 5 Stale Issues
    • 2.3. Open Issues
    • 2.4. Closed Issues
    • 2.5. Issue Discussion Insights
  • III. Pull Requests
    • 3.1. Open Pull Requests
    • 3.2. Closed Pull Requests
    • 3.3. Pull Request Discussion Insights
  • IV. Contributors
    • 4.1. Contributors

I. News

1.1 Recent Version Releases:

No recent version releases were found.

1.2 Version Information:

Please provide the version release information you would like me to analyze and summarize.

II. Issues

2.1 Top 5 Active Issues:

We consider active issues to be issues that that have been commented on most frequently within the last week. Bot comments are omitted.

  1. [BUG] [STAT:AWAITING OPENXLA-ENG] Error occurred while building TensorFlow: This issue reports a build error encountered when compiling TensorFlow on a big endian system, specifically due to a static assertion that enforces little endian architecture support in the XLA runtime code. The user highlights that while the logic appears to work on big endian systems, the assertion prevents successful compilation, and requests investigation into whether support for big endian can be added or if the assertion should be reconsidered.

    • The comments include requests for updates and additional information, a clarification of the build steps used, and a discussion about the broader challenge of supporting big endian systems in TensorFlow and XLA, noting that many parts of the codebase assume little endian and that supporting big endian would require a significant, focused effort.
    • Number of comments this week: 2
  2. [ROCm] SIGSEGV in HSA InterceptQueue when replaying a while-loop command buffer as a HIP graph (gfx950): This issue describes a segmentation fault occurring on AMD Instinct gfx950 hardware when replaying a while-loop command buffer as a HIP graph using ROCm 7.1.25442, where the fault happens during the first execution step inside the HSA runtime. The problem is isolated to the command-buffer replay path and does not occur when executing the same computation without command buffers, with additional complexity introduced by the presence of an HSA queue interceptor from rocprofiler.

    • The comments confirm that while-loop command buffers are currently unsupported in the HIP graph runtime and recommend disabling command buffers on ROCm; they also note known issues with ROCm versions prior to 7.2.4 and suggest using updated CI images, while expressing willingness to investigate further if more detailed dynamic control flow examples are provided.
    • Number of comments this week: 2
  3. [ERR: RUNTIME] XLA compilation (jit_compile=True) aborts with a fatal CHECK when compiling tf.where on the result of diagflat(tf.unique(...)): This issue describes a fatal internal assertion failure occurring during XLA compilation with jit_compile=True when compiling a TensorFlow graph that applies tf.where to the result of diagflat(tf.unique(...)). The problem does not occur in eager mode and is caused by uninitialized memory leading to nondeterministic aborts in the XLA compiler due to a failed check on dynamic dimension sizes during broadcasting.

    • The comment confirms the issue is nondeterministic and caused by uninitialized memory in dynamic dimension sizes during broadcasting, explains the root cause in the value inference step of tf2xla lowering, and references a fix that initializes these sizes to prevent the abort.
    • Number of comments this week: 1
  4. [STAT:AWAITING RESPONSE FROM CONTRIBUTOR] [ERR:PERFORMANCE] XLA: SparseTensor dense_shape diverges from eager execution for empty tf.where → tf.linalg.tensor_diag → tf.sparse.from_dense pipeline: This issue describes a discrepancy in the dense_shape of a SparseTensor when running a TensorFlow graph with XLA JIT compilation compared to eager execution, specifically involving a pipeline of tf.where, tf.linalg.tensor_diag, and tf.sparse.from_dense operations that produce empty indices and values. The problem arises because XLA's shape inference loses track of dynamic dimensions during a reshape operation, causing the dense_shape to be incorrectly reported as fully static and leading to a mismatch between compiled and eager outputs.

    • The single comment explains that the root cause is XLA dropping the dynamic dimension information during a reshape, and a proposed fix has been submitted to preserve this dynamic bit, which should correct the dense_shape output in the repro case.
    • Number of comments this week: 1
  5. [ERR:BUILD] [STAT:AWAITING RESPONSE FROM CONTRIBUTOR] XLA compilation fails with Unimplemented padding for instruction: bitcast-convert when applying tf.bitcast to the output of tf.raw_ops.Where: This issue describes a failure in XLA compilation when applying tf.bitcast to the output of tf.raw_ops.Where within a tf.function with JIT compilation enabled, despite the graph running successfully in eager mode. The error arises because the dynamic padder in XLA does not support padding for bitcast-convert instructions that change element bit width, leading to an unimplemented padding error during compilation.

    • The comment explains that the failure is due to the dynamic padder pass in XLA not handling width-changing bitcast-convert operations properly, and a fix has been proposed in a related pull request that adds support for this case along with regression and execution tests.
    • Number of comments this week: 1

2.2 Top 5 Stale Issues:

We consider stale issues to be issues that has had no activity within the last 30 days. The team should work together to get these issues resolved and closed as soon as possible.

As of our latest update, there are no stale issues for the project this week.

2.3 Open Issues

This section lists, groups, and then summarizes issues that were created within the last week in the repository.

Issues Opened This Week: 4

Summarized Issues:

  • Crashes during GPU command buffer replay and parsing errors: Multiple issues describe crashes occurring in different parts of the system, including a segmentation fault during the replay of a while-loop command buffer on AMD Instinct gfx950 hardware and a fatal CHECK failure in the HLO text parser when parsing reshape instructions with mismatched element counts. These crashes highlight stability problems in both GPU command buffer handling and HLO parsing that result in fatal errors rather than graceful failure modes.
  • [issues/46952, issues/46955]
  • LLVM code generation failures in CPU backend: The CPU backend's LLVM code generation produces LLVM IR that fails verification and causes a fatal crash when compiling fusion kernels involving bfloat16 dtype conversions. This issue indicates a lack of robustness in the LLVM IR generation path, leading to crashes instead of graceful error handling or successful compilation.
  • [issues/46954]
  • Inefficient GPU collective operation handling zero-size slices: The GPU implementation of the ragged-all-to-all collective operation inefficiently processes zero-size peer slices, causing the operation's cost to scale linearly with the number of devices despite fixed traffic. The proposed improvement is to skip zero-size sends and receives to reduce unnecessary overhead and improve performance.
  • [issues/46982]

2.4 Closed Issues

This section lists, groups, and then summarizes issues that were closed within the last week in the repository. This section also links the associated pull requests if applicable.

Issues Closed This Week: 5

Summarized Issues:

  • Build and Compilation Issues: Several issues report build failures and compilation errors related to specific platforms and plugins. These include ROCm PJRT plugin compilation errors on Ubuntu 24.04 and CUDA GPU LIT test failures due to missing binaries, which hinder successful builds and testing on targeted hardware.
  • issues/46610, issues/46627
  • Numerical and Functional Errors in Computation: There are problems with numerical correctness and function outputs after JAX compilation, such as bitwise_xor and atan2 producing NaN values on CPU and GPU but not TPU. Additionally, incorrect bias broadcast axis in cuBLASLt matrix multiplication on Nvidia GPUs leads to wrong results compared to CPU execution.
  • issues/44705, issues/46897
  • Bazel Build System Support: A request has been made to add support for Bazel modules in the XLA package to improve compatibility and adoption within Bazel builds. This enhancement aims to facilitate better integration and usage of XLA in Bazel-based environments.
  • issues/27926

2.5 Issue Discussion Insights

This section will analyze the tone and sentiment of discussions within this project's open and closed issues that occurred within the past week. It aims to identify potentially heated exchanges and to maintain a constructive project environment.

Based on our analysis, there are no instances of toxic discussions in the project's open or closed issues from the past week.


III. Pull Requests

3.1 Open Pull Requests

This section provides a summary of pull requests that were opened in the repository over the past week. The top three pull requests with the highest number of commits are highlighted as 'key' pull requests. Other pull requests are grouped based on similar characteristics for easier analysis. Up to 25 pull requests are displayed in this section, while any remaining pull requests beyond this limit are omitted for brevity.

Pull Requests Opened This Week: 23

Key Open Pull Requests

1. [XLA:FFI] Add backend-agnostic collectives FFI extension: This pull request adds a backend-agnostic extension to the XLA:FFI that exposes an XLA-owned host collective communicator to custom-call handlers, allowing custom kernel authors to reuse XLA's communicator through an opaque void* handle instead of managing their own.

  • URL: pull/46910
  • Associated Commits: 3155e, 3dc98, d27fa, 1eac3, 67d72, 75a4f

2. [GPU][Triton] Reject tiles exceeding shared memory during Triton tile selection: This pull request adds a shared-memory constraint to the Triton tile-selection process in order to reject tile sizes that exceed the device's shared memory capacity before compilation, thereby preventing RESOURCE_EXHAUSTED errors by enabling the search to fall back to smaller, compatible tiles, and includes tests covering both legacy and experimental tiling paths.

  • URL: pull/46764
  • Associated Commits: c613e, 21080, 584b3, 9c617, 16894

3. [XLA:GPU][oneAPI] Enable launching collective operations with oneCCL backend.: This pull request introduces a new feature that enables launching collective operations such as AllReduce, Broadcast, ReduceScatter, AllGather, Send/Recv, AllToAll, and CollectivePermute on Intel GPUs using the oneCCL backend within the XLA GPU oneAPI framework.

  • URL: pull/46689
  • Associated Commits: 6ae34, cdbd5, 6431c

Other Open Pull Requests

  • Triton GPU compiler all-gather backend: This pull request implements a new all-gather backend for the Triton GPU compiler using collective memory on the experimental tiling path. It extends the CollectiveKernelThunk to enable cross-device synchronization and runtime support while preserving the replica-id parameter, including end-to-end tests to validate functionality.
    • pull/46707
  • ROCm platform testing and fixes: Multiple pull requests improve ROCm support by updating upstream testing for ROCm on the therock 7.14 platform, fixing the record FFI test for ROCm by using PTX for CUDA and CUBIN for ROCm, and skipping tests unsupported on ROCm such as HostMemoryAllocateTest.Numa and ragged_all_to_all_e2e subtests. These changes enhance debugging, integration, and test stability on ROCm.
    • pull/46889, pull/46906, pull/46831, pull/46832
  • ROCm symmetric memory and communication enhancements: Several pull requests address symmetric memory support and GPU collectives on ROCm, including temporarily disabling symmetric memory support as a hot-fix with plans for version checks, implementing a GPU collectives FFI extension to reuse NCCL communication and clique management, and adding a GPU communicator FFI extension to access the XLA-owned ncclComm_t host communicator. These changes pave the way for symmetric-memory registration and improved collective operations.
    • pull/46895, pull/46911, pull/46863
  • ROCTX profiling support on ROCm: Two pull requests add and improve ROCTX profiling support for XLA annotations on ROCm platforms. They enable ScopedAnnotation to emit real ROCTX ranges integrated with profiling infrastructure and add support for capturing application-emitted ROCTX markers in the ROCm profiler timeline, enhancing profiling detail and fixing related collector defects.
    • pull/46935, pull/46933
  • XLA CPU intrinsic and symbol fixes: Pull requests fix the range predicate in the CPU log1p intrinsic to correct floating-point errors and add support for the __truncsfhf2 and __extendhfsf2 symbols on the s390x architecture to fix JIT symbol resolution failures. These changes improve numerical accuracy and kernel execution on specific CPU architectures.
    • pull/46765, pull/46823
  • ShardingParam overflow and validation improvements: This pull request addresses unguarded 32-bit integer overflow issues in the ShardingParam class by implementing overflow-safe 64-bit accumulation with explicit range checks, validating deserialized objects, and modifying ToDeviceList to fail closed to prevent errors and crashes.
    • pull/46773
  • XLA concat and dynamic padder bug fixes: Two pull requests fix bugs in XLA by modifying the HandleConcatenate function to enable concat-to-broadcast-of-reshape rewrites for binary concatenates and by adding support for bit-width-changing bitcast-convert operations in the dynamic padder. These fixes improve fusion, avoid unnecessary copies, and enable correct compilation of dynamically-sized tensors.
    • pull/46783, pull/46793
  • XLA one-sided Jacobi SVD and dynamic dimension fixes: Pull requests fix the one-sided Jacobi SVD expansion to correctly handle zero-sized matrix dimensions by returning empty singular values and identity matrices, and fix a bug in Literal::Broadcast by initializing all dynamic dimension sizes to their bounds before overwriting mapped dimensions. These changes prevent compilation failures and ensure correct behavior for dynamic shapes.
    • pull/46800, pull/46893
  • ROCm backend module caching improvement: This pull request improves the ROCm backend by caching loaded modules based on their binary content to prevent redundant module loads caused by identical binaries at different addresses, while maintaining existing reference counting and unload behavior.
    • pull/46874
  • PJRT threading and executor migration: This pull request replaces PJRT’s custom WorkerThread abstraction with the tsl::Executor API backed by single-threaded ThreadPoolAsyncWorkRunners. It migrates execution, asynchronous dispatch, and callback runners while preserving serialized execution and callback ordering, simplifying callback scheduling, and removing obsolete implementations and tests.
    • pull/46939

3.2 Closed Pull Requests

This section provides a summary of pull requests that were closed in the repository over the past week. The top three pull requests with the highest number of commits are highlighted as 'key' pull requests. Other pull requests are grouped based on similar characteristics for easier analysis. Up to 25 pull requests are displayed in this section, while any remaining pull requests beyond this limit are omitted for brevity.

Pull Requests Closed This Week: 37

Key Closed Pull Requests

1. [xla] Propagate HangWatchdog execution timeouts through CoordinationService: This pull request enhances the XLA distributed execution framework by enabling the HangWatchdog to propagate execution timeouts through the CoordinationService, allowing stalled hosts in a multi-host setup to report errors via RPC so that the central coordination service can broadcast failure notifications and trigger a clean, synchronized abort of collective operations across all nodes, thereby preventing indefinite hangs and enabling graceful job termination with clear error reporting.

  • URL: pull/44429
  • Associated Commits: 1da29, 55844, a1811, c1ab0, fca27, 8dc66, 03fd8, 51909, c0c41, b178f, a17a6, 7857f

2. [ROCm] Fix gpu_collectives_test: This pull request addresses multiple failing subtests in the gpu_collectives_test on the ROCm platform by skipping unsupported tests on non-CUDA platforms, adding missing platform communication support, implementing registered memory support for RCCL, and enabling the use of multiple clique IDs to ensure compatibility and functionality parity with the CUDA side.

  • URL: pull/46650
  • Associated Commits: af874, 044db, 58b1b, f6a94, d93af, f5acc

3. [XLA:CPU] Fix binary add post-op rank mismatch for oneDNN ndims requirement: This pull request addresses a rank mismatch issue in the oneDNN binary add post-operations for the XLA CPU backend by implementing fixes and adjustments to ensure compatibility with oneDNN's ndims requirements.

  • URL: pull/44338
  • Associated Commits: d0be3, 0f6ff, 71bac, bd3a3, 7a6cb

Other Closed Pull Requests

  • Backend support and hardware integration: Multiple pull requests focus on extending and improving backend support for various hardware platforms. These include integrating the Intel Triton XPU backend into the XLA GPU build system, adding support for the SM107 NVPTX target, and fixing Blackwell FP4 support for sm_100 and sm_120 GPU devices to enable successful Triton tests.
    [pull/43722, pull/46646, pull/46657]
  • Bug fixes and correctness improvements: Several pull requests address bugs and correctness issues in XLA components. Fixes include correcting dynamic dimension propagation in shape inference, updating buffer assignment to prevent unsafe overlaps, rerouting DSFv2 rewriter users to remove duplicated operations, and fixing device description strings for consistency.
    [pull/46543, pull/46678, pull/45799, pull/46339]
  • Optimization and compiler enhancements: Some pull requests introduce new optimizations and compiler improvements. These include algebraic simplifications for xor operations to ensure IEEE 754 compliance, adding typed collective communication domains for better scheduling, and enabling pre-order and post-order traversal hooks in the ThunkSequence API for enhanced analysis.
    [pull/46594, pull/46541, pull/46749]
  • Library and dependency updates: Updates to external libraries and dependencies are proposed to maintain compatibility and improve functionality. This includes upgrading the oneDNN library to v3.12.3 for the CPU backend and adding missing runfiles for ROCm profiling libraries to fix runtime linking issues.
    [pull/46133, pull/46777]
  • GPU backend and CUDA toolchain improvements: Enhancements to the GPU backend and CUDA toolchain support include enabling PTX version 9 selection for CUDA 13 toolchains, adding cuDNN graph API lowering for non-GEMM pointwise fusions, and disabling certain backends when binary libraries are disabled to fix failing tests.
    [pull/46647, pull/46703, pull/44378]
  • Profiling and tracing fixes: Fixes related to profiling and tracing on ROCm platforms improve accuracy and clarity. These include correcting stream_id tracking by using HIP stream handles and adding dense stream remapping for better trace viewer lane naming.
    [pull/46182]
  • Code quality and testing infrastructure: Improvements to code quality and testing include adding OSS-Fuzz libFuzzer harnesses for HLO text-format parser and proto deserialization, and introducing a pre-commit configuration for clang-format, clang-tidy, and buildifier checks to improve developer efficiency.
    [pull/42055, pull/46417]
  • Custom call and FFI improvements: Enhancements to the PJRT GPU custom call extension and FFI error handling standardize and improve flexibility. These include adding a traits field for custom call handlers and unifying FFI error handling with an ErrorPolicy abstraction layer.
    [pull/45377, pull/46801]
  • Collective operations and replica group handling: Improvements to collective operations include adding a flag to bypass profitability heuristics in all_reduce_splitter optimizations and conditionally wrapping fusion parameters into replica-id pointer tables to fix memory-access faults in GPU collectives.
    [pull/46521, pull/46352]

3.3 Pull Request Discussion Insights

This section will analyze the tone and sentiment of discussions within this project's open and closed pull requests that occurred within the past week. It aims to identify potentially heated exchanges and to maintain a constructive project environment.

Based on our analysis, there are no instances of toxic discussions in the project's open or closed pull requests from the past week.


IV. Contributors

4.1 Contributors

Active Contributors:

We consider an active contributor in this project to be any contributor who has made at least 1 commit, opened at least 1 issue, created at least 1 pull request, or made more than 2 comments in the last month.

If there are more than 10 active contributors, the list is truncated to the top 10 based on contribution metrics for better clarity.

Contributor Commits Pull Requests Issues Comments
shawnwang18 83 1 0 0
ezhulenev 28 9 0 22
alekstheod 28 1 0 0
mfrancepillois 23 3 1 1
draganmladjenovic 9 3 0 13
kodlan 11 8 0 3
pemeliya 11 5 0 2
sohaibiftikhar 0 0 0 17
kanvi-nervana 14 1 0 0
akhilgoe 14 0 0 0

Don't miss what's next. Subscribe to Weekly Project News:
Powered by Buttondown, the easiest way to start and grow your newsletter.