Weekly Project News

Archives
Subscribe

Weekly GitHub Report for Pytorch: September 21, 2026 - September 28, 2026 (20:37:56)

Weekly GitHub Report for Pytorch

Thank you for subscribing to our weekly newsletter! Each week, we deliver a comprehensive summary of your GitHub project's latest activity right to your inbox, including an overview of your project's issues, pull requests, contributors, and commit activity.


Table of Contents

  • I. News
    • 1.1. Recent Version Releases
    • 1.2. Other Noteworthy Updates
  • II. Issues
    • 2.1. Top 5 Active Issues
    • 2.2. Top 5 Stale Issues
    • 2.3. Open Issues
    • 2.4. Closed Issues
    • 2.5. Issue Discussion Insights
  • III. Pull Requests
    • 3.1. Open Pull Requests
    • 3.2. Closed Pull Requests
    • 3.3. Pull Request Discussion Insights
  • IV. Contributors
    • 4.1. Contributors

I. News

1.1 Recent Version Releases:

The current version of this repository is v2.6.0

1.2 Version Information:

Released on January 29, 2025, PyTorch 2.6 introduces significant enhancements including torch.compile support for Python 3.13, the new performance tuning API torch.compiler.set_stance, and improved AOTInductor packaging and ABI compatibility. Notable highlights also include FP16 support on X86 CPUs, expanded Intel GPU support, a backward-incompatible security change flipping the default of torch.load to weights_only=True, and the deprecation of official Conda packages, reflecting a trend toward improved performance, security, and streamlined deployment.

II. Issues

2.1 Top 5 Active Issues:

We consider active issues to be issues that that have been commented on most frequently within the last week. Bot comments are omitted.

  1. [MODULE: PERFORMANCE] [MODULE: BUILD] [TRIAGED] [ENHANCEMENT] [MODULE: DEBUG-BUILD] [BOT-TRIAGED] Editable install: DEBUG=1 builds stall ~200s at Making editable step, deflating a wheel that pip immediately unpacks: This issue addresses a significant performance bottleneck during the pip install -e . editable install step for PyTorch when built with DEBUG=1, where the build process stalls for around 200 seconds due to single-threaded zlib compression of large debug shared libraries into a wheel that pip immediately unpacks and discards. The problem stems from scikit-build-core's mandatory DEFLATE compression of editable wheels, and the discussion explores potential fixes including disabling compression for editable installs or using a persistent install tree to avoid the costly compression and unpacking cycle, with upstream improvements pending.

    • The comments confirm the issue's reproducibility and discuss trade-offs between compression speed and rebuild overhead, suggest workarounds using scikit-build-core's existing options, report a related artifact location bug that was fixed upstream, and provide timelines for when the upstream fixes might be released and integrated.
    • Number of comments this week: 7
  2. [MODULE: DOCS] [TRIAGED] [MODULE: DOC INFRA] [docs] Docs website search finds duplicates and produces bad snippets: This issue addresses the problem of the PyTorch documentation website's search functionality returning multiple duplicate links for the same function and producing poor-quality snippets that do not effectively summarize the content. It highlights the need for improved search result deduplication, better snippet generation, and enhanced relevance ranking, especially for exact symbol name queries, to improve the overall user experience when searching the docs.

    • The comments discuss various examples of duplicate search results and poor snippets, suggest potential solutions like post-processing to deduplicate and rerank results, mention ongoing experiments with different search engines, and note partial fixes on the main branch while emphasizing the need for further improvements in snippet quality and search relevance.
    • Number of comments this week: 6
  3. [MODULE: BUILD] [TRIAGED] [MODULE: REGRESSION] [MODULE: XPU] [BOT-TRIAGED] [XPU] Building Pytorch XPU with older ccache (4.9 - 4.11) fails: This issue reports a build failure of PyTorch XPU on Ubuntu 24.04 when using older versions of ccache (4.9 to 4.11) due to improper handling of SYCL compiler flags, specifically the stripping of -Xarch_host options causing linkage errors. The problem is traced to a commit that enabled ccache support for SYCL kernels compilation, which requires ccache version 4.12 or higher for proper DPC++ compiler support and flag preservation.

    • The comments discuss testing various ccache versions confirming the issue occurs with versions below 4.12 and is resolved in 4.12 and above. They deliberate on two possible fixes: disabling ccache support for SYCL compilation in PyTorch’s top-level CMakeLists or applying a version check in the torch-xpu-ops build scripts, with a preference expressed for the latter approach.
    • Number of comments this week: 5
  4. [HIGH PRIORITY] [TRIAGE REVIEW] [MODULE: ERROR CHECKING] [TRIAGED] [MODULE: CORRECTNESS (SILENT)] [MODULE: MPS] [BOT-TRIAGED] [MPS] float32 reduction over a permuted bf16 view larger than 231 bytes silently returns zeros**: This issue describes a problem on Apple MPS where performing a float32 reduction over a large permuted bf16 tensor view silently returns zeros when the memory in flight exceeds the recommended maximum working set size, rather than raising an error or producing correct results. The root cause is identified as a non-contiguous-source cast operation losing its output when the combined command buffers reference more memory than the device's recommended limit, with a workaround involving synchronizing operations to avoid this silent failure.

    • The comments detail attempts to reproduce the issue, clarifications that the failure is due to memory in flight rather than tensor size, a self-sizing reproducer script, and confirmation that the problem persists in nightly builds unless synchronization is used; a fix to surface out-of-memory errors more clearly is also discussed.
    • Number of comments this week: 5
  5. [TRIAGE REVIEW] [MODULE: PERFORMANCE] [MODULE: CPU] [MODULE: REGRESSION] [OP-BENCH] [MODULE: LINEAR ALGEBRA] [MODULE: ARM] [BOT-TRIAGED] [RELEASE TRIAGE] [NEEDS REPRODUCTION] [TRIAGED] aarch64 CPU op-benchmark: matmul 256×256 trans_b ~25× slower than baseline: This issue reports a significant performance regression in the aarch64 CPU operator benchmark for a specific 256×256 matrix multiplication configuration with the second matrix transposed, which runs approximately 25 times slower than the expected baseline. The user seeks confirmation on whether the expected performance should be around 234 ms or closer to 6,000 ms to determine if this is a genuine regression in the GEMM/transpose path on the main branch.

    • The comments show an initial willingness to tackle the issue, discussion of a strategy to reproduce and diagnose the regression, requests for workflow approvals to run benchmarks, and identification of a likely thread affinity problem with a proposed fix that improved performance in testing; there is also a request for insight into how the root cause was determined to aid learning.
    • Number of comments this week: 4

2.2 Top 5 Stale Issues:

We consider stale issues to be issues that has had no activity within the last 30 days. The team should work together to get these issues resolved and closed as soon as possible.

As of our latest update, there are no stale issues for the project this week.

2.3 Open Issues

This section lists, groups, and then summarizes issues that were created within the last week in the repository.

Issues Opened This Week: 144

Summarized Issues:

  • Incorrect numerical results and silent failures in automatic differentiation and compiled operations: Multiple issues report silent incorrect outputs or numerical inaccuracies in PyTorch's forward-mode and higher-order automatic differentiation, Inductor backend, and torch.compile operations. These include errors in derivative computations, in-place mutations causing wrong results, dtype and shape inconsistencies in compiled NumPy functions, and incorrect handling of symbolic integers and views, all leading to silent failures without errors or warnings.
    • issues/198017, issues/198031, issues/198033, issues/198052, issues/198056, issues/198064, issues/198068, issues/198070, issues/198071, issues/198081, issues/198094, issues/198095, issues/198100, issues/198101, issues/198102, issues/198118, issues/198119, issues/198131, issues/198155, issues/198187, issues/198270, issues/198280, issues/198307, issues/198332, issues/198343, issues/198364, issues/198382, issues/198422, issues/198424, issues/198484, issues/198485, issues/198486, issues/198533, issues/198545, issues/198553, issues/198555, issues/198567, issues/198568, issues/198669, issues/198676, issues/198868, issues/198851, issues/198852, issues/198853
  • Performance regressions and optimization issues on specific hardware and backends: Several issues highlight performance degradations and suboptimal kernel implementations on Apple M4 GPUs, Blackwell GPUs, Metal backend, and others. These include slower FlexAttention throughput, missing NVFP4 inference tactics, excessive kernel launches due to output-scale kernels, and inefficient native Metal matmul kernels, all impacting training and inference speed.
    • issues/198037, issues/198126, issues/198171, issues/198172, issues/198173, issues/198174, issues/198175, issues/198182, issues/198128, issues/198375
  • Random number generation and numerical accuracy issues in compiled and vectorized code: Multiple reports describe accuracy losses and incorrect RNG behavior in Inductor and compiled backends, including erfinv and erf accuracy degradation near boundaries, float64 RNGs producing float32 precision values, float16/bfloat16 RNGs generating out-of-range values, and inconsistent dropout RNG states during checkpointing, causing silent numerical errors and incorrect gradients.
    • issues/198133, issues/198183, issues/198856, issues/198862, issues/198863, issues/198866
  • Bugs and crashes related to inductor backend and compilation pipeline: Several issues report crashes, assertion failures, and compilation errors in Inductor and AOTInductor backends due to vectorization bugs, incorrect lowering of operations, memory corruption, and scheduler errors. These include heap corruption, C++ compile errors, crashes on Windows with OpenMP, and failures in lowering permute or foreach_pow operations.
    • issues/198332, issues/198613, issues/198624, issues/198205, issues/198422, issues/198450, issues/198503, issues/198880
  • Discrepancies and bugs in dtype handling and tensor layout between eager and compiled modes: Multiple issues describe inconsistent dtype promotions, incorrect dtype inference, and layout mismatches between eager execution and compiled code, causing silent errors or crashes. Examples include torch.any returning bool instead of uint8, torch.full dtype inference errors, torch.cat dropping dtype promotion, and stride mismatches on transposed inputs.
    • issues/198056, issues/198081, issues/198485, issues/198486, issues/198364, issues/198851, issues/198852, issues/198693, issues/198694
  • Bugs in PyTorch's autograd, Dynamo, and torch.compile control flow and exception handling: Several issues report incorrect behavior in compiled code related to exception handling, generator reconstruction, dictionary mutation, and try/except blocks, causing silent failures or crashes that differ from eager execution. These include exceptions escaping compiled calls, incorrect generator state after graph breaks, and failed dictionary mutations in compiled functions.
    • issues/198070, issues/198189, issues/198190, issues/198192, issues/198197
  • Incorrect or inconsistent behavior in PyTorch distributions and statistical functions: Multiple issues describe numerical inaccuracies, incorrect gradients, or NaN outputs in distribution log_prob, cdf, icdf, and mode calculations, as well as precision problems in special functions like polygamma and bessel functions, affecting probabilistic modeling and statistical computations.
    • issues/198578, issues/198579, issues/198580, issues/198582, issues/198583, issues/198663
  • Memory management, serialization, and IPC bugs causing crashes or incorrect behavior: Several issues report memory corruption, heap corruption, serialization bugs that change in-place operations to out-of-place, and IPC cache misses due to uninitialized padding, leading to crashes or silent incorrect results.
    • issues/198077, issues/198305, issues/198372, issues/198794, issues/198609
  • Platform-specific bugs and build issues on Windows, ROCm, XPU, and AArch64: Multiple issues describe crashes, build failures, or incorrect behavior on specific platforms including Windows OpenMP unload crashes, ROCm SVD driver bugs, XPU cache crashes, and AArch64 duplicate symbol problems, impacting platform stability and compatibility.
    • issues/198205, issues/198425, issues/198432, issues/198499, issues/198522, issues/198869
  • Test failures and disabled tests on XPU and ROCm platforms: Several tests have been disabled due to consistent failures on XPU and ROCm platforms, indicating ongoing stability or compatibility issues in these environments.
    • issues/198180, issues/198181, issues/198709, issues/198712, issues/198713, issues/198714, issues/198445
  • Requests for documentation, feature support, and clarifications: Some issues request improved documentation for device module APIs, support for new reduction types in scatter operations, clarification on NativeRT usage, and addition of new loss functions implemented in C++ for stability and efficiency.
    • issues/198222, issues/198566, issues/198657, issues/198667
  • Miscellaneous bugs including optimizer state restoration, SSL certificate expiry, and setuptools version limits: Additional issues include optimizer state dict ignoring parameter names causing silent mismatches, SSL certificate expiry warnings, and build breakage due to setuptools version upper limits in rolling release environments.
    • issues/198693, issues/198694, issues/198847, issues/198444

2.4 Closed Issues

This section lists, groups, and then summarizes issues that were closed within the last week in the repository. This section also links the associated pull requests if applicable.

Issues Closed This Week: 121

Summarized Issues:

  • Segmentation faults and crashes in tensor operations: Multiple issues report segmentation faults and crashes occurring in various tensor operations such as linalg.ldl_solve, fractional_max_pool3d, torch.lu_unpack, torch.ge, torch.atan2, and torch.igamma. These crashes are often caused by improper input usage, dimension mismatches, or unsupported tensor types, and many have been resolved in newer PyTorch versions by replacing crashes with proper error handling.
    • issues/91590, issues/91633, issues/92776, issues/92794, issues/92818, issues/92828, issues/193823
  • torch.compile related errors and regressions: Several issues describe problems with torch.compile, including runtime errors with DeepSpeed ZeRO Level 3 sharding, extremely long compile times with the inductor backend and dynamic=True, failures with non-scalar outputs in autograd.grad, unsupported argument types like set, and bugs causing incorrect tensor mutations or skipped bounds checks. These issues cause compilation failures, incorrect outputs, or performance regressions.
    • issues/115484, issues/133734, issues/145899, issues/139603, issues/197829, issues/195665
  • Memory safety and buffer overflow vulnerabilities: Issues report heap-buffer-overflow and out-of-bounds memory writes in functions like torch.nn.utils.rnn.pack_padded_sequence with quantized tensors, inductor::_alloc_from_pool creating out-of-bounds tensor views, and MPS backend operations (index_copy_, index_fill_, index_reduce_, embedding_bag) failing to validate indices properly, leading to memory corruption and potential crashes.
    • issues/114928, issues/194053, issues/189968, issues/189969, issues/189970, issues/189971
  • CUDA and GPU backend errors and performance regressions: Multiple issues describe CUDA runtime errors such as invalid configuration arguments in torch.cdist, missing Python headers causing build failures, performance regressions in Inductor's convolution backward layout optimization, and training accuracy regressions on NVIDIA H100 GPUs due to cuDNN and kernel changes. These problems affect runtime stability and training performance on GPU platforms.
    • issues/128791, issues/196977, issues/189239, issues/195483
  • Documentation and API inconsistencies: Some issues highlight missing or incorrect documentation, such as undefined ProcessGroupOptions references, unclear ordering contracts for unpack_hook calls, and inconsistent argument validation in functions like torch.histc between CPU and CUDA backends. These inconsistencies cause confusion and potential misuse of APIs.
    • issues/130958, issues/198014, issues/196008
  • Test failures and disabled tests on specific platforms: Several tests are disabled or fail consistently on platforms such as ROCm, XPU, and CUDA due to unsupported features, flaky behavior, or hardware incompatibilities. This includes tests for CUTLASS on AMD GPUs, redispatch tests on XPU, and fused attention tests on XPU.
    • issues/168801, issues/168847, issues/168848, issues/168849, issues/168861, issues/196748, issues/197334, issues/197335, issues/197336, issues/197337, issues/198319
  • Export and serialization issues: Problems occur during model export and deserialization, including failures to handle data-dependent expressions in torch.distributions.Normal, inability to handle dynamic sequence lengths in scaled dot product attention, and a critical security vulnerability in torch.export.load() due to unvalidated code injection from placeholder tensor names.
    • issues/134533, issues/135061, issues/170127, issues/198506
  • Dynamo and TorchDynamo related bugs: Issues include guard recompilation crashes due to NoneType subscripting, incorrect handling of static methods causing AttributeErrors, and problems with recorded execution strategies overriding skip directives, leading to crashes or incorrect tracing behavior.
    • issues/176211, issues/188544, issues/196520
  • MPS backend functional and performance issues: The MPS backend suffers from missing implementations for fractional max pooling and segment reduce, incorrect variance calculations in reductions, loss of precision in torch.expm1, and severe performance regressions in torch.histc and torch.linalg.solve_triangular. These issues degrade functionality and performance on Apple Silicon devices.
    • issues/197847, issues/197880, issues/197235, issues/198304, issues/198708
  • Type errors and symbolic integer handling: Several issues report cryptic or uninformative error messages related to symbolic integer arrays, type mismatches in operator registration, and bugs in torch.compile with isinstance checks involving runtime-checkable Protocols, causing confusion and silent failures during compilation or runtime.
    • issues/140960, issues/163501, issues/195969
  • Floating point exceptions (SIGFPE) in tensor operations: Multiple operations such as avg_pool1d with large kernel sizes, native_channel_shuffle with zero groups, and unfold_backward with specific inputs cause floating point exceptions leading to crashes.
    • issues/145065, issues/146787, issues/146791
  • Distributed and process group issues: Problems include inability to initialize distributed process groups on CPU devices due to incorrect device_id requirements, and the need for a method to abort all process groups concurrently when a rank exits prematurely to prevent hangs.
    • issues/160731, issues/173815
  • Cache and logging verbosity problems: Excessive and overly detailed cache logging output from TORCH_LOGS=+inductor and failures to initialize TORCH_LOGS and TORCH_TRACE environment variables cause noisy logs and hinder debugging.
    • issues/137875, issues/121318
  • Compiler and fusion pass bugs in Inductor: Bugs include crashes during fusion passes like merge_select_cat_aten, race conditions causing non-deterministic gradients in fused reduction kernels, and import errors escaping during NVGEMM availability probes, leading to compilation failures or incorrect results.
    • issues/198559, issues/198562, issues/197463
  • Incorrect or inconsistent behavior between eager and compiled modes: Discrepancies include torch.clamp silently ignoring unrepresentable bounds in compiled mode, adaptive_max_pool2d accepting int64 inputs in compiled mode but raising errors in eager mode, and differences in exception string formatting inside compiled functions.
    • issues/195675, issues/197062, issues/197110
  • Security vulnerabilities: A critical security issue exists in torch.export.load() where unvalidated placeholder tensor names are injected into dynamically generated Python code, enabling attacker-controlled code execution during deserialization before any validation occurs.
    • issues/198506
  • Build and CI infrastructure problems: Issues include missing files in release tarballs causing build failures, Git clone failures due to certificate verification errors, and long queuing delays in CI runners caused by node offline states and permission misconfigurations.
    • issues/152532, issues/198420, issues/198076

2.5 Issue Discussion Insights

This section will analyze the tone and sentiment of discussions within this project's open and closed issues that occurred within the past week. It aims to identify potentially heated exchanges and to maintain a constructive project environment.

Based on our analysis, there are no instances of toxic discussions in the project's open or closed issues from the past week.


III. Pull Requests

3.1 Open Pull Requests

This section provides a summary of pull requests that were opened in the repository over the past week. The top three pull requests with the highest number of commits are highlighted as 'key' pull requests. Other pull requests are grouped based on similar characteristics for easier analysis. Up to 25 pull requests are displayed in this section, while any remaining pull requests beyond this limit are omitted for brevity.

Pull Requests Opened This Week: 402

Key Open Pull Requests

1. [POC] cpp symnode: This pull request introduces a proof-of-concept implementation of a native C++ symbolic node ("SymNode") system in PyTorch, featuring core expression handling, assumption management, symbolic operations, native guards, and integration with Python bindings to enable efficient symbolic computation and evaluation within the framework.

  • URL: pull/198653
  • Associated Commits: 3478a, a7ab6, 9d72a, 018bd, ac3a8, e056e, 11d32, 9a13c, 94ad5, a8a6c, 69d55, 36c9f, 799a3, 303c1, 13431, 5f819, c9d7e, 771fd, d9a35, fd029, 35d4c, a28ea, b6be8, d0326, f4241, 648cd, 348c6, c38d4, 13f3d, 62cf7, 9e921, 2c326, b5c49, e92b3, b9a4e, 811f9, 51256, 6ecdd, 10ba5, 73cef, 7410d, 601b0, af428, a0d25, 441b4, 42709, b6ca2, dcbbd, 3618c, 88604, 1dc6a, 8cc5a, 31573

2. [precompile] Make DynamoTracer the default tracer and settle the docs: This pull request changes the default tracer in the capture() API from MakeFxTracer to DynamoTracer, enabling multi-call tracing by default, updates the documentation to reflect this change, and adjusts tests accordingly to ensure that DynamoTracer is the default tracer while preserving MakeFxTracer as an explicit opt-in for single-call or lambda-based captures.

  • URL: pull/198217
  • Associated Commits: 7a1bf, 22bb0, a13bc, 7c437, 63e4f, 0c840, b8bf9, fdad4, f6fba, 8765a, b52e7, e9c6e, 239e6, 6a0ad, 0a30f, 9c65e, c5546, 46aec, 830ac, 0a124, 51117, 49059, ab7b8, 9b6ae, 38d3f, c9d7a, 11d68, fe879, a43dc, f46e8, b1bda, bbb38, 88751, c4155, 04f72, 9f54a, b828f

3. [precompile] Add DynamoTracer and serve its captures through capture(): This pull request introduces DynamoTracer and _DynamoCapture to extend torch.compiler.precompile.capture() from supporting only single-call MakeFxTracer captures to multi-call, graph-breaking, and keyword-argument-capable captures that record and serialize multiple execution variants into a standalone artifact served in a fresh process, enabling robust precompilation of dynamic PyTorch functions including training steps with gradient accumulation, while adding concurrency safety, detailed error gating, and new tracer configuration knobs.

  • URL: pull/198138
  • Associated Commits: ec884, 9f054, 7066e, b1490, 57419, 8c3e6, ca15f, 7a2dd, a2243, 339bd, 17b55, a1714, d1a8a, cd99f, 0167d, 3f255, 21716, b56c2, f52a5, fff36, 098b2, 483b5, a93fa, 442be, 917d4, 93e8e, 29fe2, fd3c1, b093c, 87153, 620ab, 6a773, 7b60e, fd43d, 5f350, 9248d

Other Open Pull Requests

  • Dynamo multi-call capture and checkpointing: Multiple pull requests enhance PyTorch Dynamo's multi-call capture capabilities by introducing PrecompileSession for managing multi-graph captures, enabling rendering of compilation results as standalone artifacts, and adding checkpointing and summary methods to _DynamoCapture. These changes improve artifact serialization, guard management, and recovery during tracing, supporting caller-driven mid-block checkpointing and faithful reloads of multi-call captures.
  • [pull/198136, pull/198137, pull/198214, pull/198215, pull/198216, pull/198227]
  • Precompile capture robustness and concurrency: Several pull requests improve the robustness and concurrency safety of the precompile capture process by adding a process-wide reentrant lock to serialize captures, ensuring RNG state is preserved during capture, and validating tensor-subclass input dtype/device at load time. These enhancements prevent race conditions, unintended RNG side effects, and silent errors, enabling safe nested precompile calls and consistent artifact loading.
  • [pull/198245, pull/198246, pull/198247]
  • export_python decorator improvements: A series of pull requests refine the torch.compiler.export_python decorator by adding features such as capturing compiled Python source to a file for later execution, pinning machine-type constraints for artifacts, normalizing function signatures to reject unsafe arguments, restoring RNG state consistency, introducing detailed PrecompileErrors for artifact issues, and serializing concurrent captures with a reentrant lock. These changes enhance usability, error handling, reproducibility, and concurrency safety of the export_python workflow.
  • [pull/198252, pull/198254, pull/198255, pull/198256, pull/198252, pull/198253]
  • FSDP test and training optimizations: Improvements to Fully Sharded Data Parallel (FSDP) include skipping reference training when device errors are expected, removing unnecessary delays in tests, skipping unsupported parameter combinations early, and deferring gradient upcasting to reduce GPU time and memory overhead during mixed-precision training. These optimizations speed up tests and improve training efficiency without changing FSDP behavior or correctness.
  • [pull/198066, pull/198082]
  • Distributed and backend test efficiency: Test runtime is significantly reduced by reusing supervisors across distributed test suites and reusing worker processes in backend collective tests, along with reorganizing test helpers and separating tests requiring fresh processes. These changes maintain test correctness while improving execution speed.
  • [pull/198059, pull/198130]
  • Support for aggregate types in Triton kernels: Foundational support is added in Dynamo and Inductor for capturing, flattening, and handling recursively nested aggregate types like tuples and NamedTuples as FX-literal metadata for user-defined Triton kernels. This includes structured argument reconstruction, validation, autotuning, and epilogue fusion, preparing for enhanced Python wrapper code generation and kernel specialization.
  • [pull/198532, pull/198534]
  • CUDA architecture modeling in AOT compilation: CUDA architecture families are explicitly modeled in PyTorch's native AOT compilation process, enabling declarative CUDA compatibility, improved target selection preferring widest compatible architectures, and fallback to JIT for unsupported devices. This enhances CUDA support and compilation flexibility.
  • [pull/198229]
  • NVSHMEM dependency update and compatibility fixes: PyTorch's NVSHMEM dependency is updated from version 3.7.2 to 3.8.0, including changes to CUDA CI/binary build images and pip dependencies for CUDA 13 wheels. Compatibility issues in Triton SHMEM tests are fixed by dynamically exposing updated NVSHMEM signal operation enums, maintaining backward compatibility without source changes.
  • [pull/198406]

3.2 Closed Pull Requests

This section provides a summary of pull requests that were closed in the repository over the past week. The top three pull requests with the highest number of commits are highlighted as 'key' pull requests. Other pull requests are grouped based on similar characteristics for easier analysis. Up to 25 pull requests are displayed in this section, while any remaining pull requests beyond this limit are omitted for brevity.

Pull Requests Closed This Week: 751

Key Closed Pull Requests

1. Dynamo 3.15 support: This pull request proposes adding support for Dynamo version 3.15 to the PyTorch project, although it has not been merged.

  • URL: pull/178393
  • Associated Commits: 6f92f, fd536, ea7a9, 984dd, e2687, eadba, d09db, 2d316, 41d0c, 3511c, 11d7e, 65a6c, 4448c, 3c994, 2b9ce, bf28d, 7b1d3, 68f1e, 883f8, b4b7b, 84670, 380f4, 7b924, 71935, a0e06, b78f8, c3448, 4d6dd, b09a6, 90f71, 37214, b0e8e, c9a42, 60b33, a1e8c, 813f0, e44ca, 465c7, 43a53, 50664, 8ec25, 4a7f4, 74b04, 46adb, 446ed, a03a5, 1b6e8, 2d1a7, 7d90a, 4fccf, e6645, 11129, 6190a, b0713, 24650, 7929d, 5a0b5, c18f6, 5bb21, 57807, 24d66, cb6a3, 96d8b, d277c, 648ac, 42be8, 9c8ae, 82bd7, 965cb, 32d58, 2c34d, 00ec1

2. Add docker image for python 3.15: This pull request proposes adding a Docker image for Python 3.15 to the PyTorch project, but it was not merged.

  • URL: pull/179094
  • Associated Commits: 49830, cad99, 470ff, 1b441, 13706, 52d82, daec0, 4a781, e0c08, e488f, c0d99, 284d7, caea1, bc92a, 9dd54, 9811b, 2a6f5, 9cd49, 642bd, 6255a, 0b4d2, ea6e9, ed7ac, 536c3, 410c7, 82873, f627b, d5460, b4152, 400e3, 59303, 8c104, a881a, 52a1b, 4ca63, 34ed6, a0664, f2350, 9874e, 37430, e2c4e, b139c, e4da0, e66ff, 64cda, 6c9ed, ea67b, 29045, 69147, 39037, d5155, 61916, ee560, c5661, 45038, f3f36, d5db8, ab5fe, 7b809, e9870, 0c1d8, a0eaf, ff701, dcb9f, 2739c, a5dec, eec26, ab67c, a3123

3. [no-ci] [precompile] Pin the module surface and the training= grad-mode contract: This pull request pins the public API surface of torch.compiler.precompile and the training-mode gradient contract by explicitly asserting module and type hint consistency, directly tests that the captured call runs under the grad mode dictated by the training= argument (overriding ambient mode), and adds comprehensive tests covering API re-homing, file artifact mode preservation, autocast behavior during capture and serve phases, and torch_function mode application, all without changing implementation behavior but strengthening correctness guarantees and forward compatibility.

  • URL: pull/197344
  • Associated Commits: 9db8d, dc604, e0950, 2c7e7, a7ed1, 57d44, 621ce, f6969, 8c2ea, 0d8af, 69672, 2b2ca, fede7, e9d86, b3b60, 5b725, 1b26d, 2b40a, 0241b, 7ede0, 0885e, 5c8ac, 38a71, 8ecc1, 4508c, 42652, 941a9, 5ae40, bae70, 8a2a9, fea06, 264e3, 87868, d71a3, 8ccb0, df9ec, 75ca8, a3118, 5f836, afbf1, 43e81, b4c16, 013c1, 59d4b, ac2d4

Other Closed Pull Requests

  • Precompile API Refactoring and Enhancements: Multiple pull requests refactor the torch.compiler.precompile module from a callable singleton to a proper module exposing a public API with capture(), load(), Capture, and MakeFxTracer. These changes include switching to file-based artifact paths, adding caller-driven capture/load functions with atomic writes and error handling, porting tests to the new API, and tagging artifacts with tracer and serving mode metadata to improve integration and validation.
  • [pull/197343, pull/197341, pull/197387, pull/197388, pull/197456]
  • DynamoTracer Contract Tests and Serving Integration: Pull requests reintroduce DynamoTracer contract tests as expected failures to preserve tracer behavior validation and implement serving of DynamoTracer captures through the capture() function. This makes DynamoTracer the default tracer, enables atomic saving/loading of captured artifacts as PrecompiledModules with keyword argument support, and improves serialization and reporting while passing most contract tests.
  • [pull/197347, pull/197353]
  • PrecompileSession Multi-Graph Capture and Snapshotting: A pull request adds a snapshot_artifact() method to PrecompileSession that renders multi-graph captures as standalone artifacts including python code, serialized backend data, and a standalone driver. This enables repeated, self-contained execution of captured variants with detailed guard checks and support for inductor subgraphs, while handling unsupported frames and tensor default arguments.
  • [pull/197352]
  • Guard Invariants Reporting and Diagnostic Improvements: Several pull requests enhance guard invariant reporting during precompilation captures by tracking guard completeness and safety, and improve diagnostic processes for unpicklable values in the PyTorch Dynamo guard state. These include capturing scope roots before pruning, extending search to entire guard state structures and function internals, and appending precise attribute paths to unpicklable values to improve error reporting and identification.
  • [pull/197351, pull/197761, pull/197760, pull/197758, pull/197759]
  • Profiling and CUPTI Integration Enhancements: Multiple pull requests introduce CUPTI monitor benchmarks for profiling performance, add an in-process NCCL metadata plugin, propose communication monitoring tests and plugins for CUPTI integration, and introduce CUDA-graph correlation utilities. These contributions improve profiling capabilities, decode speed, export efficiency, and communication monitoring in PyTorch.
  • [pull/186655, pull/187518, pull/187520, pull/187519, pull/187517]
  • Serialization and Round-Tripping of Dynamo Eager Graphs: A pull request improves serialization of Dynamo eager graphs by changing EagerCacheArtifact to carry generated source code and module state instead of re-deriving graphs on load. This enables lossless round-tripping of graphs including nodes without Proxy targets and introduces mechanisms to rebuild artifacts as executable source modules while preserving nested graph bodies as pickled blobs.
  • [pull/197349]
  • Memory-Aware Fusion Guard in Inductor Scheduler: A pull request introduces a configurable memory timeline peak-aware fusion guard that estimates and restricts fusion candidates which would increase peak memory usage beyond a threshold or extend large output buffers. This prevents fusions that are locally beneficial but globally detrimental to memory efficiency in PyTorch's Inductor scheduler.
  • [pull/190344]
  • Namespace Module Names for Inductor Python Modules: A pull request introduces the namespace_module_names function to rewrite and suffix top-level definitions in multiple Inductor Python modules. This allows multiple compiled subgraphs to coexist and be executed independently within a single Python environment by avoiding name collisions.
  • [pull/197348]
  • Autocast Neutralization in Precompiled Artifacts: A pull request implements a mechanism to neutralize ambient autocast effects in precompiled PyTorch artifacts by wrapping served calls in a context that disables autocast dispatch keys. This ensures output dtypes match those baked into the capture rather than being altered by surrounding autocast regions during execution.
  • [pull/197455]
  • Link Time Optimization (LTO) Attempts: A pull request attempts to enable LTO in the open-source PyTorch by addressing related bugs, disabling LTO on Windows, narrowing IPO support, and fixing issues discovered through LTO testing. However, this attempt was ultimately not merged.
  • [pull/180186]
  • s390x ZVECTOR Data Representation Rework: A pull request reworks the s390x ZVECTOR data representation by replacing emulation of 32-byte vectors with direct support for 16-byte vectors, enabling ZVECTOR compilation in CI, updating vectorized operations, and fixing related tests.
  • [pull/173072]

3.3 Pull Request Discussion Insights

This section will analyze the tone and sentiment of discussions within this project's open and closed pull requests that occurred within the past week. It aims to identify potentially heated exchanges and to maintain a constructive project environment.

Based on our analysis, there are no instances of toxic discussions in the project's open or closed pull requests from the past week.


IV. Contributors

4.1 Contributors

Active Contributors:

We consider an active contributor in this project to be any contributor who has made at least 1 commit, opened at least 1 issue, created at least 1 pull request, or made more than 2 comments in the last month.

If there are more than 10 active contributors, the list is truncated to the top 10 based on contribution metrics for better clarity.

Contributor Commits Pull Requests Issues Comments
bobrenjc93 4751 335 3 17
d4l3k 355 43 0 17
liangel-02 218 16 0 6
jananisriram 211 8 0 4
malfet 154 20 3 45
fffrog 2 0 0 184
guangyey 153 8 0 18
Native-Neo 165 1 0 0
weifengpy 129 12 0 10
drisspg 134 4 0 11

Don't miss what's next. Subscribe to Weekly Project News:
Powered by Buttondown, the easiest way to start and grow your newsletter.