Weekly GitHub Report for Kubernetes: September 21, 2026 - September 28, 2026 (20:38:38)
Weekly GitHub Report for Kubernetes
Thank you for subscribing to our weekly newsletter! Each week, we deliver a comprehensive summary of your GitHub project's latest activity right to your inbox, including an overview of your project's issues, pull requests, contributors, and commit activity.
Table of Contents
I. News
1.1 Recent Version Releases:
The current version of this repository is v1.32.3
1.2 Version Information:
The Kubernetes version released on March 11, 2025, introduces key updates detailed in the official CHANGELOG, with additional binary downloads available. For comprehensive information on new features and changes, users are encouraged to refer to the Kubernetes announce forum and the linked CHANGELOG.
II. Issues
2.1 Top 5 Active Issues:
We consider active issues to be issues that that have been commented on most frequently within the last week. Bot comments are omitted.
-
[PRIORITY/IMPORTANT-SOON] [SIG/NODE] [KIND/FLAKE] [TRIAGE/ACCEPTED] [Flaking Test] E2eNode Suite [It] [sig-node] [DRA] [Feature:DynamicResourceAllocation] Two resource Kubelet Plugins must call NodeUnprepareResources again if it's in progress for one plugin when Kubelet restarts [Serial]: This issue concerns a flaky end-to-end test in the Kubernetes node dynamic resource allocation feature, where two Kubelet plugins must call NodeUnprepareResources again if one plugin's call is in progress when the Kubelet restarts. The test fails due to a timing issue where the Kubelet restarts while a container creation is still in progress, causing a long timeout on stopping the pod sandbox and delaying the NodeUnprepareResources call beyond the test's wait period.
- The comments discuss the cause of the test flake related to a test rename affecting Testgrid display, detailed analysis of the failure sequence involving container creation and pod sandbox stop delays, and a proposed fix to wait for the blocked call before restarting the Kubelet. The fix is under review, and related issues in the CRI-O runtime are being tracked separately.
- Number of comments this week: 12
-
[SIG/NODE] [KIND/FLAKE] [SIG/TESTING] [PRIORITY/IMPORTANT-LONGTERM] [TRIAGE/ACCEPTED] [WG/DEVICE-MANAGEMENT] [Flaking Test] [sig-node] pull-kubernetes-kind-dra-all API timeouts: This issue describes intermittent API timeouts and client connection losses occurring in the pull-kubernetes-kind-dra-all presubmit job, likely caused by CPU throttling under the job's 2 CPU quota which stalls the job and leads to failures in cleanup steps and log checks. The investigation includes local replays showing heavy CPU throttling at 2 CPUs and successful test runs with increased CPU limits, prompting a proposed increase of the CPU quota to 4 to mitigate the flake.
- The comments discuss diagnosing the flake as CPU throttling due to the 2 CPU limit, share local replay results confirming heavy throttling and improved outcomes with more CPUs, and coordinate updating the job configuration to increase CPU resources to 4 CPUs to resolve the issue.
- Number of comments this week: 10
-
[SIG/SCALABILITY] [KIND/FAILING-TEST] [NEEDS-TRIAGE] [SIG/K8S-INFRA] [Failing test] ci-kubernetes-e2e-kops-aws-500-node-dra-with-workload-amazonvpc-using-cl2: control-plane instance not created: This issue describes a failure in the Kubernetes CI job for a 500-node cluster on AWS where the control-plane instance was not created, causing the cluster creation to never complete and subsequent tests to fail due to inability to connect to the API server. The root cause was identified as a misconfiguration in the kOps kubetest2 deployer, which searched for a deprecated role name and skipped necessary control-plane overrides, and a fix was implemented and merged to address this problem.
- The comments detail the identification of the root cause related to the role naming mismatch in the deployer, the submission and review of a fix in the kOps project, acknowledgments from maintainers, and a follow-up test-infra update to use multiple control-plane instance types to improve job stability, with ongoing verification of the fix in CI.
- Number of comments this week: 8
-
[PRIORITY/IMPORTANT-SOON] [SIG/NODE] [TRIAGE/ACCEPTED] Kubelet emits spurious Pod status updates due to nondeterministic allocated resource health ordering: This issue describes a problem in kubelet v1.36.3 where the order of entries in Pod status updates related to allocated resource health can change nondeterministically, causing unnecessary Pod status updates and watch event churn despite no actual health changes. The root cause is the iteration over maps without a stable order when constructing resource status slices, and the proposed solution is to canonicalize the ordering of these resource status entries to ensure stable output across repeated syncs.
- The comments include a reference to an earlier test fix for non-determinism, expressions of intent to investigate the issue, assignment and prioritization of the issue, and a prompt for the original reporter to submit a fix, indicating active triage and planned resolution steps.
- Number of comments this week: 7
-
[PRIORITY/BACKLOG] [SIG/NODE] [KIND/FEATURE] [TRIAGE/ACCEPTED] Leverage remote partitions for exclusive CPUs (Linux 6.7+): This issue proposes migrating the Kubernetes CPU manager to leverage the new "remote partitions" feature introduced in Linux 6.7, which allows containers to allocate exclusive CPUs without modifying other cgroups, thereby eliminating the current reconcile loop and improving CPU allocation consistency and isolation. The proposal outlines writing assignable CPUs to
cpuset.cpus.exclusiveacross the cgroup hierarchy and setting partitions on container creation or resize, aiming to reduce race conditions, protect system processes from CPU contention, and simplify CPU management.- The comments include questions clarifying the cgroup hierarchy and partition ownership, the role of the default CPU set after removing the reconcile loop, and isolation expectations for host processes; contributors express support for the proposal, discuss the need for coordination with other components like NRI/OCI, consider the implications of moving CPU actuation to SyncPod, and highlight challenges around concurrency and retry mechanisms in the new design.
- Number of comments this week: 7
2.2 Top 5 Stale Issues:
We consider stale issues to be issues that has had no activity within the last 30 days. The team should work together to get these issues resolved and closed as soon as possible.
As of our latest update, there are no stale issues for the project this week.
2.3 Open Issues
This section lists, groups, and then summarizes issues that were created within the last week in the repository.
Issues Opened This Week: 40
Summarized Issues:
- Pod and Kubelet Status Issues: Several bugs cause incorrect pod readiness and status reporting, including nondeterministic Pod status updates causing churn, pods incorrectly marked Ready during startup probe failures, and race conditions causing MirrorPod restart status flapping. These issues lead to inconsistent pod states and unreliable status information in Kubernetes clusters.
- [issues/142293, issues/142305, issues/142422]
- Flaky and Failing Tests: Multiple flaky tests affect Kubernetes e2e and node test suites, including intermittent failures in NFS volume execution, CustomResourceValidationRules, ResourceQuota lifecycle, volume attach/detach reconciler, and node resource plugin unprepare behavior. These flakes cause unreliable test results and complicate CI stability.
- [issues/142308, issues/142321, issues/142324, issues/142349, issues/142404]
- Informer Cache Staleness and Controller Bugs: Informer cache lag leads to controllers acting on stale data, causing premature pod deletions, volume detachments, and job scheduling failures due to stale PVC or node taint information. This stale data breaks safety guarantees and blocks job completions in gang-scheduled jobs.
- [issues/142330, issues/142331, issues/142332]
- API Conversion and Admission Webhook Issues: Problems with API conversion tools and admission webhook configurations include missing method implementations causing unversioned old objects, stale ValidatingWebhookConfiguration caches causing admission failures, and conversion-gen tool failing on cross-package types. These issues disrupt API versioning and webhook reliability.
- [issues/142302, issues/142345, issues/142374]
- Resource Management and CPU Allocation Bugs: Bugs in resource resizing and CPU allocation cause potential oversubscription, out-of-memory errors, and race conditions where CPUs are prematurely reassigned during pod downsizing. These issues highlight the need for improved resource management and CPU tracking in the kubelet.
- [issues/142361, issues/142363]
- Hugepages Validation Bug: The AllowIndivisibleHugePagesValues flag incorrectly exempts all hugepages entries from divisibility checks if any single value is indivisible, allowing invalid hugepages configurations to pass validation. This undermines pod validation correctness for hugepages resource requests.
- [issues/142357]
- Dependency and Library Maintenance Concerns: The Kubernetes project depends on unmaintained libraries like
gopkg.in/inf.v0and maintains a custom CEL extension library instead of using the upstream CEL project, raising concerns about bus factor risks and missing out on performance improvements. Proposals include switching to upstream CEL and adopting upcoming performance enhancements. - [issues/142364, issues/142366, issues/142367]
- Documentation and Deprecation Updates: There are stale documentation links needing updates, plans to deprecate older AuthorizationConfiguration versions, and additions of Windows end-to-end tests with updated documentation on supported stop signals. These efforts aim to keep documentation accurate and Kubernetes APIs current.
- [issues/142318, issues/142365, issues/142370]
- Metrics and Histogram Inconsistencies: Issues with histogram metrics include unset NativeHistogramMinResetDuration, kube-apiserver exposing classic instead of native histograms due to feature gate application order, and cloud-controller-manager failing to properly enable native histograms despite reporting them enabled. These cause inconsistent Prometheus metrics across components.
- [issues/142377, issues/142379, issues/142454]
- Kubectl and CLI Tooling Bugs:
kubectl describe podfails to show projected volume source details for certain types, andkubectl execsilently truncates stdout output when the client reads slower than the pod writes, leading to incomplete data without errors. These bugs reduce CLI usability and visibility into pod state. - [issues/142309, issues/142376]
- Cluster Setup and Resource Exhaustion Failures: Cluster bring-up fails due to image URL newline issues from multiple COS images in the same family, and GCP resource exhaustion prevents provisioning of large VM types in CI jobs, causing cluster setup failures and blocking tests. These infrastructure issues impact cluster reliability and test execution.
- [issues/142362, issues/142423]
- Load Balancing and Client Connection Improvements: A proposal to add naive load balancing support in client-go for multiple connections to the same endpoint aims to improve traffic distribution and reliability when interacting with kube-apiserver replicas. This enhancement targets client performance and fault tolerance.
- [issues/142383]
- Resource Plugin and ImageGC Test Configuration: Suggestions include making ImageGCPeriod and other wait durations configurable in e2e_node tests to reduce long test times, addressing timing issues in resource plugin unprepare calls during kubelet restarts, and fixing permission bugs on volume subpath directories after kubelet restarts. These changes improve test efficiency and runtime correctness.
- [issues/142417, issues/142404, issues/142464]
- CPU Manager Migration Proposal: A proposal to migrate the Kubernetes CPU manager to use Linux 6.7 remote partitions aims to eliminate reconcile loops, improve CPU assignment consistency, and simplify resource updates by leveraging new kernel features for exclusive CPU allocation. This would enhance CPU isolation and management.
- [issues/142430]
- Race Conditions and Data Races in Kubelet: A data race in kubelet's use of cAdvisor's GetVfsStats causes test failures due to concurrent writes, requiring an upstream fix and library update. This race condition affects kubelet stability under race-enabled testing.
- [issues/142439]
- Dependency Removal Tracking: Tracking the removal of unwanted grpc dependencies reintroduced in grpc v1.84.0 and fixed upstream in v1.86.0 is ongoing to ensure Kubernetes dependencies remain clean and up to date.
- [issues/142483]
2.4 Closed Issues
This section lists, groups, and then summarizes issues that were closed within the last week in the repository. This section also links the associated pull requests if applicable.
Issues Closed This Week: 32
Summarized Issues:
- Event resource validation errors: Setting the
eventTimefield in KubernetesEventresources fails due to strict validation requiring additional fields likereportingController,reportingInstance, andaction. This prevents the resource from being accepted by the API, causing issues in event creation or updates.
- Informer synchronization in tests: There is a need to consistently use
WaitForCacheSyncafter starting sharedInformerFactory in unit and integration tests to ensure proper informer synchronization. This practice helps coordinate multiple SIG groups and track related pull requests effectively.
- Race conditions causing stale connection tracking: A race condition causes UDP connection tracking entries to remain in Conntrack tables during pod termination because kube-proxy only cleans up entries when the pod reports
Serving=false. This leads to stale entries persisting and client request failures during pod rotation.
- Scheduler blocking due to synchronous queue processing: The Kubernetes scheduler's
moveAllToActiveOrBackoffQueuefunction blocks operations and causes scheduler hangs under high-frequency pod updates. Adding support for asynchronous execution of this function is requested to prevent these blocking issues.
- OIDC authentication error message clarity: Using the email claim for OIDC authentication results in an unhelpful "invalid bearer token" error that hides the root cause related to the
email_verifiedclaim being false. This makes diagnosing and tracing authentication failures difficult.
- Incorrect CPU request reporting in Pod status: The CPU request value in a Pod's
status.resourcesis incorrectly lowered after kubelet updates due to inaccurate conversions between CPU shares and weights. This causes the reported CPU request to become invalid compared to the originally specified millicores.
- Performance bug in OpportunisticBatching scheduling: A redundant synchronous call to
RunFilterPluginsin the OpportunisticBatching feature causes unnecessary overhead and prevents effective batching for low-resource pods. Redesigning the rescore logic eliminated this overhead and improved batching efficiency.
- Pod creation allowed despite resource request-limit mismatch: Pods are successfully created with the PodLevelResources feature enabled even when pod-level resource requests exceed pod-level limits, contrary to expectations that such pods should be blocked or error out.
- Flaky test due to resource slice generation race: The
TestControllerSyncPoolsuite intermittently fails due to multiple resource slices existing with different resource pool generations. This flaky behavior may be linked to a race condition involving the MutationCache.
- Sensitive manifest URL headers logged in cleartext: When kubelet is configured with
--manifest-urland credential-bearing--manifest-url-header, these sensitive headers are logged in cleartext at startup. This bypasses redaction mechanisms and leads to secret disclosure of upstream-manifest credentials.
- Windows e2e Security Context test failure: The Kubernetes e2e Security Context test fails on Windows nodes because it uses a Linux-specific pod running the
id -ucommand, which is unavailable in Windows containers. This causes the pod to fail to start and time out.
- Inconsistent pool generation completeness due to input order: The
GatherPoolsfunction incorrectly determines pool generation completeness based on the order of input ResourceSlices when theirResourceSliceCountvalues disagree. This leads to inconsistent allocation behavior depending on input permutation.
- CPU Manager static policy socket affinity bug: Enabling both
distribute-cpus-across-cores=trueandalign-by-socket=truecauses CPUs to be allocated across multiple sockets instead of being confined to a single socket. This violates the intended socket affinity guarantee.
- Device-taint eviction controller misses DRA-backed pods: The device-taint eviction controller fails to evict pods using Dynamic Resource Allocation (DRA)-backed extended-resource claims because it only checks
pod.Spec.ResourceClaims. Pods using the classic extended-resource request path are evicted correctly, but DRA pods remain running despite NoExecute taints.
- Data race in resource.Quantity during read-only operations: The
resource.Quantitytype is unsafely mutated during read-only operations likeString()andSize(), causing concurrent access conflicts. This data race primarily affects kubelet volume allocation and pod management code paths.
- gRPC-Go heap exhaustion vulnerability (CVE-2026-84304): A security vulnerability in gRPC-Go causes heap memory exhaustion via HTTP/2 DATA frame fragmentation. Discussions focus on updating dependencies to patched versions while managing unwanted dependencies after a failing CVE scan.
- Windows node devicePath field undefined semantics: The
Node.status.volumesAttached[].devicePathfield has undefined semantics and platform-specific code issues on Windows nodes. A proposal suggests redefining it to carry the CSIVolumeHandleand updating API documentation while fixing dead Unix-specific code.
- WindowsHostNetwork feature cleanup: The WindowsHostNetwork feature is being fully cleaned up by removing its feature gate constant, special-case test bypasses, and dead test code following its deprecation and withdrawal from Kubernetes.
- DRA kubelet plugin socket path length bug: The DRA kubelet plugin helper can generate rolling-update service socket paths exceeding the maximum allowed length for AF_UNIX pathname endpoints when using long driver names. This results in unusable automatic service socket endpoints during rolling updates.
- CVE placeholders for future details: Placeholders exist for CVE-2026-2270 and CVE-2026-76654, with details to be populated after their public releases.
- nftables localhost nodeport proxy starts unnecessarily: The nftables localhost nodeport proxy starts on 127.0.0.1 even when the nodeport address does not include 127.0.0.1. The proxy should only start if the nodeport address explicitly contains 127.0.0.1 or its subnet.
- Job controller race condition on delete and recreate: Deleting and immediately recreating a Job with the same name causes the job controller to get stuck tracking the old Job's status. This prevents the new Job from progressing and creating pods.
- kubectl describe node resource usage display bugs: The
kubectl describe nodeoutput has two defects: per-pod resource usage percentages show large negative numbers on nodes without allocatable resources due to division by zero, and resource usage rows disagree with totals during in-place pod resizing because of inconsistent use of status versus spec resources.
- Kubelet /containerLogs endpoint returns misleading HTTP 200: When a container is removed but its Pod still exists, the kubelet's
/containerLogsendpoint returns HTTP 200 with an error message in the body. This causes clients likekubectl logsto misinterpret the failure as a successful response.
- Flaky pod logs retrieval over websockets: Pods intermittently fail to support log retrieval over websockets due to URL scheme rewriting issues. Retries dial plaintext websocket URLs against TLS ports, resulting in bad status errors.
- kubectl accepts empty resource names in resource lists: Kubectl commands incorrectly accept resource list entries with empty resource names, leading to ResourceQuota objects with empty-string keys. Validation is proposed to reject empty resource names for consistent parsing.
- Flaky ci-kubernetes-e2e-capz-master-windows test due to timeouts: The ci-kubernetes-e2e-capz-master-windows test flakes daily due to timeout errors waiting for machine deployments to reach the desired state, despite no failed AzureMachines being detected.
- Patch release needed for CVE-2026-39821 in Go 1.35: A patch release of Kubernetes version 1.35 is needed to fix CVE-2026-39821, which affects version 1.35.8 despite being resolved in Go 1.26.6 or later.
- Node Declared Features framework extension for pod re-admission: A proposal extends the Node Declared Features framework to allow features to differentiate between first pod admission and re-admission after kubelet restart. This enables opting out of strict re-admission enforcement to avoid unnecessary pod eviction when features are disabled but non-critical.
- DRA E2E test failure due to disabled feature detection: The ResourceSlice controller helper fails to detect disabled features, causing dropped ResourceSlice fields when running against clusters without dynamic resource allocation enabled. This leads to DRA E2E test failures.
- Kubernetes GPU device plugin e2e test flaking: The GPU device plugin e2e test flakes repeatedly due to timeouts running
nvidia-smiandcuda-demo-suitecommands on AWS and GCE providers. This causes pods to fail to start or complete within expected time.
2.5 Issue Discussion Insights
This section will analyze the tone and sentiment of discussions within this project's open and closed issues that occurred within the past week. It aims to identify potentially heated exchanges and to maintain a constructive project environment.
Based on our analysis, there are no instances of toxic discussions in the project's open or closed issues from the past week.
III. Pull Requests
3.1 Open Pull Requests
This section provides a summary of pull requests that were opened in the repository over the past week. The top three pull requests with the highest number of commits are highlighted as 'key' pull requests. Other pull requests are grouped based on similar characteristics for easier analysis. Up to 25 pull requests are displayed in this section, while any remaining pull requests beyond this limit are omitted for brevity.
Pull Requests Opened This Week: 105
Key Open Pull Requests
1. WIP: DRA: allocator performance optimizations: This pull request implements multiple layered performance optimizations to the Device Resource Allocation (DRA) scheduler in Kubernetes, improving scheduling throughput significantly by reducing redundant conversions and enabling lazy evaluation of device attributes and capacities, while also fixing correctness bugs and supporting new features like CompositePodGroup, thereby enhancing both current performance and future extensibility of device allocation.
- URL: pull/142406
- Associated Commits: b7600, aaa29, 0d328, 624bf, 75a76, 3e282, 59fb1, aa886, c134c, 73192, 7b204, 3ad42, 6ee47, 578b9, 0ff0e, 64eef, 73c54, a074b, bac19, 1c2bb
2. Migrate apimachinery JSON encoding to json/v2: This pull request migrates the Kubernetes apimachinery JSON encoding from the original encoding/json package to the newer json/v2 package while preserving v1 semantics, improving performance and memory usage across various realistic workloads, and includes comprehensive tests and benchmarks to ensure compatibility and efficiency.
- URL: pull/142359
3. [WIP] Move the apiserver signal handlers into a stdlib-only leaf package: This pull request refactors the Kubernetes apiserver by moving the signal handling functions SetupSignalHandler, SetupSignalContext, and RequestShutdown from the generic k8s.io/apiserver/pkg/server package into a new, standard-library-only leaf package k8s.io/apiserver/pkg/server/signals, thereby reducing dependencies and import graph complexity for components like kubelet and kube-proxy, which no longer need to import the generic API server or the CEL stack, resulting in smaller binary sizes and cleaner import boundaries.
- URL: pull/142453
Other Open Pull Requests
- Job controller pod expectation fixes: Multiple automated cherry picks fix a bug in the job controller that clears pending Pod expectations when deleting a Job. This prevents issues where a newly recreated Job with the same name remains unreconciled due to inherited expectations from the previous Job.
- [pull/142390, pull/142391, pull/142392]
- Kubelet container startup status fix: A bug fix in the kubelet ensures that containers still within their startup period are not incorrectly reported as started after a kubelet restart when the ChangeContainerStatusOnKubeletRestart feature gate is disabled. This is achieved by adding a check for the pod spec's startup probe in the isContainerStarted function to ensure accurate container status reporting.
- [pull/142303]
- Hugepages pod validation improvement: Pod validation for hugepages values is fixed by replacing a broad skip flag with a precise per-value exact match ratchet. This ensures that only unchanged stored indivisible hugepages values bypass validation, while new or different values are fully validated.
- [pull/142358]
- ML-DSA token support introduction: Work-in-progress support for ML-DSA tokens is introduced, including vendor updates and new test cases for the OIDC authenticator. This adds foundational support for ML-DSA tokens in Kubernetes authentication.
- [pull/142397]
- DeviceTaintRule deviceSelector 'all' field: A new boolean field
allis added toDeviceTaintRule.spec.deviceSelectorto explicitly match all devices across every driver in the cluster. This removes reliance on an empty selector, suppresses related warnings, and enforces mutual exclusivity with other selector fields. - [pull/142287]
- Proxier HTTP proxy logging and tests: Detailed verbose logging is added to the
NewProxierWithNoProxyCIDRfunction to improve HTTP proxy decision visibility for egress auditing and debugging. The PR also includes focused test coverage for IPv4 and IPv6 CIDR scenarios and renames a probe HTTP source file without changing functionality. - [pull/142288]
- Preemption-aware scoring and fixes: An initial proof of concept for preemption-aware scoring is implemented, including fixes for PodInfo resource cache race conditions and updates to partial preemption tests. The SignPlugin on DefaultPreemption is implemented with bypassing of batching during preemption.
- [pull/142325]
- Pod group runtime hierarchy validation: Validation logic is added to the kube-scheduler to enforce constraints on pod group runtime hierarchies, such as maximum tree depth, consistent Workload references, gang parent/child group type restrictions, and disruption mode compatibility. These constraints follow the CompositePodGroup KEP specifications.
- [pull/142375]
- HTTP/2 configuration update in apiserver: The Kubernetes apiserver and apimachinery are updated to configure HTTP/2 using the standard net/http package instead of the deprecated golang.org/x/net/http2 functions. Related tests are refactored, compatibility with Go 1.27 is ensured, and direct imports of the deprecated http2 package are removed.
- [pull/142432]
- DRADeviceBindingConditions GA graduation: The
DRADeviceBindingConditionsfeature gate is graduated from Beta to GA in Kubernetes v1.38 by locking it enabled by default. The stable DRA allocator is updated to supportDeviceBindingAndStatus, obsolete beta API comments are removed, and related tests are adjusted accordingly. - [pull/142478]
- kubectl certificate caching optimization: kubectl performance is improved by reducing redundant file I/O operations related to reading certificate data. Proper caching mechanisms are implemented to reuse certificate information efficiently, benefiting systems affected by on-access virus scanning.
- [pull/142286]
- TAS performance test parametrization: The number of nodes per rack in TAS performance tests is parametrized to enable multi-rack placements. This enhances topology aware scheduling evaluation by generating multiple candidate placements for podgroup scheduling.
- [pull/142316]
- Validating admission performance improvement: Performance is improved by skipping re-decoding of managedFields in the validating admission wrapper when managedFields remain unchanged. This results in significant reductions in processing time, memory allocations, and allocations per operation.
- [pull/142320]
- Volume recycler timeout bug fix: A bug fix clamps the Kubernetes volume recycler timeout value at
math.MaxInt32and adds a guard to prevent integer overflow during multiplication. This ensures recycler pods for very large persistent volumes can be created without exceeding API server limits. - [pull/142334]
- Scheduler inter-pod affinity caching: A caching mechanism is implemented for the scheduler's inter-pod affinity and topology spread plugins to optimize CPU-intensive computations. Results are memoized per (namespace, pod name, node) with incremental invalidation, rebasing previous work and fixing related test defects.
- [pull/142335]
- Kubelet resource health status ordering: The ordering of allocated resource health statuses in the kubelet is stabilized by canonicalizing them by resource name and resource ID. This ensures deterministic serialization and prevents unnecessary Pod status updates when resource health remains unchanged.
- [pull/142340]
- etcd3 backend prefix normalization: Handling of backend prefix slashes in etcd3 storage is simplified by normalizing the backend prefix to exclude a trailing slash. This allows direct concatenation with resource keys starting with a slash, improving code clarity without user-facing changes.
- [pull/142347]
- GvkParser duplicate GVK tolerance: The GvkParser in apimachinery is improved to tolerate duplicate GroupVersionKind entries by retaining the first occurrence and skipping duplicates. This fixes a bug causing UnstructuredExtractor construction failures with overlapping OpenAPI definitions and includes regression tests.
- [pull/142388]
- terminationGracePeriodSeconds API doc fixes: The API documentation for
terminationGracePeriodSecondsis fixed and clarified regarding pod deletion behavior and kubelet termination handling, including preStop hook timing. Obsolete wording related to theProbeTerminationGracePeriodfeature gate is removed, and inaccuracies about zero value interpretation are corrected without runtime changes. - [pull/142446]
- ParseQuantity zero value parsing fix: A bug fix allows
ParseQuantityto correctly parse zero values with a minimum int32 exponent (e.g., "0e2147483648") without error while still rejecting nonzero values with such exponents. This restores compatibility with version 1.37.1 behavior and improves test coverage for fractional exponent parsing. - [pull/142452]
- Kubelet gRPC probe client update: The kubelet's gRPC probe client is updated to use the non-deprecated
grpc.NewClientinstead ofgrpc.DialContextwithWithBlock. This results in immediate failure on connection refusal rather than retrying until timeout, changes failure messages for unresponsive servers, and fixes container lifecycle tests to use valid exit codes. - [pull/142457]
- IPAddress resource validation graduation: The IPAddress resource's four declarative validation rules are graduated from beta to stable by removing beta wrappers and handwritten validation checks. This enforces required and immutable constraints on
parentRefinIPAddressSpecand required fields inParentReferenceacrossnetworking/v1andnetworking/v1beta1API versions. - [pull/142473]
- Watchcache selection predicate support: Watchcache functionality is enhanced by adding support for selection predicates in watch validation and validating the correctness of the GetList operation.
- [pull/142474]
3.2 Closed Pull Requests
This section provides a summary of pull requests that were closed in the repository over the past week. The top three pull requests with the highest number of commits are highlighted as 'key' pull requests. Other pull requests are grouped based on similar characteristics for easier analysis. Up to 25 pull requests are displayed in this section, while any remaining pull requests beyond this limit are omitted for brevity.
Pull Requests Closed This Week: 128
Key Closed Pull Requests
1. KEP-5958: Add server-side opt-out for metadata.managedFields: This pull request introduces an alpha feature gate ManagedFieldsOptOut that allows clients to opt out of receiving metadata.managedFields in API responses by specifying drop=metadata.managedFields in the Accept header, improving performance and reducing payload size across all verbs, resource types, and supported media formats in the Kubernetes API server.
- URL: pull/139561
2. kubelet: log StaticPodURLHeader key names only, not values: This pull request fixes a bug in the kubelet where sensitive credential values from the StaticPodURLHeader, such as Authorization tokens, were being logged in plaintext at startup by modifying the logging behavior to only record the header key names instead of their values, thereby preventing secret leakage in runtime logs and adding regression tests to ensure this protection.
- URL: pull/140220
3. apimachinery/reflect: add DeepEqualWithOptions to support option IgnoreUnexportedFields: This pull request introduces a new function DeepEqualWithOptions in the apimachinery reflection package that enhances the existing deep equality check by allowing options such as ignoring unexported fields, enabling safer and more flexible comparisons of Kubernetes API objects with in-memory representations, particularly benefiting use cases involving structs with unexported fields like those used by CRD users such as Istio and Cilium.
- URL: pull/126998
Other Closed Pull Requests
- Apiserver Cacher Cleanup: This pull request cleans up the apiserver cacher by removing the OrderedListPrefix and fully typing the store interface to use *Element instead of interface{}, eliminating unnecessary boxing and unboxing. These changes simplify the read and write paths and improve type safety without changing any behavior.
- OOM Watcher for Windows: This pull request adds a Windows-specific implementation to the kubelet OOM watcher that periodically reconciles CRI container statuses to detect containers terminated by out-of-memory conditions. It emits Kubernetes OOMKilled events and increments the container_oom_events_total metric, enabling OOM observability on Windows nodes where kernel cgroup OOM streaming is unavailable.
- Quantity API Testing and Validation Enhancements: This pull request adds comprehensive test cases covering edge cases and behavioral changes in the Quantity API and related validation functions, ensuring no regressions in resource quantity handling. The tests were informed and constructed using AI-assisted analysis to address issue #141166.
- Development Tools Update: This pull request updates several development tools, including Mockery, modfile, and protoc-gen-go, to their latest versions using the project's update scripts. These updates ensure the tooling is current without introducing any user-facing changes.
- Namespace Declarative Validation: This pull request enables declarative validation for Kubernetes namespaces by focusing on validating the
metadata.namefield through specialized validation functions. This improves how namespace resources are wired for validation in the API machinery.
- CompositePodGroup Scheduling Enhancements: This pull request enables the
MinGroupCountfield inCompositePodGroup.Spec.SchedulingPolicy.Gangto be mutable across API versionsv1alpha3andv1beta1, allowing dynamic resizing of gang scheduling requirements. It also implements a scheduler plugin enhancement that re-enqueues pods whenMinGroupCountdecreases to improve scheduling responsiveness.
- Pod Group Runtime Hierarchy Validation: This pull request is a proof of concept adding status updates, validation logic, and error exposure for pod groups and composite pod groups within the kube-scheduler. It aims to validate pod group runtime hierarchies.
- Memory Manager Drift Tolerance Configuration: This pull request introduces the
memoryManagerPolicyOptionsfield to KubeletConfiguration, enabling configuration of thememory-drift-toleranceoption under theMemoryManagerDriftTolerancefeature gate. This allows the kubelet's static memory manager to tolerate bounded per-NUMA-node memory drift on restart, improving stability by avoiding kubelet start failures due to minor kernel memory layout changes.
- Preemption Logic for PodGroups: This work-in-progress proof of concept proposes adjustments to the preemption triggering logic for subsequent scheduling attempts. It extends the scheduling decision tree to handle both flat PodGroups and CompositePodGroups with different policies to prioritize binding or preemption based on scheduling state and group hierarchy.
- Apiserver 404 JSON Response: This pull request updates the Kubernetes apiserver to return a JSON-formatted Status object with a 404 status code and appropriate Content-Type for requests to unregistered or unknown
/apis/...paths. This aligns error responses with the rest of the API instead of returning a plain-text "404 page not found" response.
- Scheduler Device Resource Allocation Bug Fix: This pull request fixes a bug in the Kubernetes scheduler's structured device resource allocator by preventing over-committing shared counters of devices dropped from ResourceSlices. It ensures the allocator rejects new counter-consuming allocations when devices are missing and adjusts handling of admin-access claims to avoid double-counting.
- Fractional Byte Warning Fix: This pull request fixes false fractional-byte warnings by avoiding overflow issues when detecting fractional byte quantities larger than the int64 range. It ensures accurate validation of resource requests and limits for Pods and PersistentVolumeClaims.
- ML-DSA Certificate Support: This pull request updates CertificateSigningRequest validation and kube-controller-manager signing flags to support ML-DSA private keys. It enforces specific usage constraints on ML-DSA certificate requests and adds tests to ensure proper issuance of ML-DSA signed certificates.
- Dynamic Resource Allocation Test Fix: This pull request adds an option to skip cgroup limit and pod status checks in the
WaitForPodResizeActuationfunction during DRA node allocatable resize tests. This addresses test failures by relying on explicit verification of DRA-inflated cgroups and pod status elsewhere in the test suite.
- Resource.Quantity Concurrency Safety: This pull request makes the String method and encoding operations of resource.Quantity read-only to prevent data races caused by mutations during concurrent use. It caches the string representation upon parsing to eliminate unpredictable mutations when handling non-canonical input or protobuf encoding.
- Validation Generator Duration Support: This pull request enhances the validation generator to support
time.Durationfields by allowing+k8s:minimumand+k8s:maximumtags to accept quoted Go duration strings. It ensures generated code interprets these bounds in nanoseconds, produces user-friendly error messages, and rejects integer payloads for such fields.
- Kubelet Device Manager Bug Fix: This pull request fixes a bug in the kubelet device manager by preserving in-flight device reservations across allocatedDevices rebuilds. This prevents the same device ID from being assigned simultaneously to multiple pods.
- Kubelet Cgroup Test Coverage: This pull request expands unit test coverage for kubelet's cgroup update and read calls to improve test reliability. It supports the removal of cgroup coverage from end-to-end tests.
- Storagemigration API Linter Enablement: This pull request enables the
commentstartkube-api-linter rule for the storagemigration API group by fixing 12 godoc comment violations in v1 and v1beta1 types. It removes storagemigration from linter exceptions and regenerates related swagger, protobuf, openapi, and applyconfiguration files.
- Device Resource Allocation Validation Bug Fix: This pull request addresses a bug in DRA validation logic by ensuring device capacity is re-validated when the
allowMultipleAllocationsflag changes from true to false or is unset. This prevents acceptance of invalidrequestPolicyconfigurations that were previously bypassed due to an optimization skipping validation if the capacity map was unchanged.
- PodSpec activeDeadlineSeconds Declarative Validation: This pull request migrates the
activeDeadlineSecondsfield inPodSpecto use declarative validation tags as part of KEP-5073. It replaces handwritten Go validation with declarative rules registered in shadow mode and updates existing validation logic accordingly.
- SchedulingAlgorithm Refactor: This pull request refines the internal
SchedulingAlgorithmtype by exporting key entry points for node selection and pod placement without requiring aScheduler. It splits node fitting logic into distinct phases and consolidates common reservation preparation code to improve structure and usability.
- Test Cancellation Refactor: This pull request removes redundant explicit calls to tCtx.Cancel in tests by relying on automatic cancellation of TContext created by ktesting.Init. It refactors deferred cancellations and teardowns to use t.Cleanup/tCtx.Cleanup callbacks to preserve proper shutdown ordering and prevent race conditions or panics.
- Kubeadm ML-DSA Encryption Support: This pull request adds support for ML-DSA encryption algorithms in kubeadm by allowing "ML-DSA-44", "ML-DSA-65", and "ML-DSA-87" as valid ClusterConfiguration.EncryptionAlgorithm values. It also corrects the keyEncipherment bit setting to apply only to RSA keys.
- Proxy Localhost NodePort Bug Fix: This pull request fixes a bug in the proxy component by refining detection of explicit localhost NodePort addresses. It ensures only 127.0.0.1 (IPv4) and ::1 (IPv6) are treated as localhost, preventing other loopback addresses like 127.0.0.2 from being incorrectly identified, and includes regression tests.
3.3 Pull Request Discussion Insights
This section will analyze the tone and sentiment of discussions within this project's open and closed pull requests that occurred within the past week. It aims to identify potentially heated exchanges and to maintain a constructive project environment.
Based on our analysis, there are no instances of toxic discussions in the project's open or closed pull requests from the past week.
IV. Contributors
4.1 Contributors
Active Contributors:
We consider an active contributor in this project to be any contributor who has made at least 1 commit, opened at least 1 issue, created at least 1 pull request, or made more than 2 comments in the last month.
If there are more than 10 active contributors, the list is truncated to the top 10 based on contribution metrics for better clarity.
| Contributor | Commits | Pull Requests | Issues | Comments |
|---|---|---|---|---|
| pohly | 56 | 6 | 3 | 42 |
| thc1006 | 48 | 9 | 4 | 40 |
| dims | 37 | 22 | 0 | 6 |
| serathius | 28 | 19 | 1 | 13 |
| jdzikowski | 18 | 5 | 1 | 35 |
| antcybersec | 22 | 13 | 2 | 12 |
| jpbetz | 18 | 2 | 0 | 27 |
| yongruilin | 33 | 1 | 0 | 7 |
| danwinship | 27 | 2 | 0 | 11 |
| liggitt | 16 | 0 | 0 | 22 |