Resolved -
We are marking this incident as resolved.
The measures already put into place have had the desired effect on the service and have improved performance and stability.
While this incident is marked as resolved, we are continuing executing on our action plan to improve the performance and reliability of our Managed Kubernetes Control Planes. Our current focus:
- Further load balancing on our Control Plane Clusters
- Migration of control planes to improved infrastructure
Aug 10, 19:58 UTC
Update -
Migration of customer workloads on the first control plane cluster has been completed ahead of schedule. Customer clusters have been redistributed across expanded infrastructure and all migrated workloads are running without issues. We continue to observe stable control plane performance with no new disruptions reported.
Aug 7, 05:06 UTC
Monitoring -
Memory limit adjustments have been fully rolled out across all control plane clusters and are showing positive effects. We have significantly expanded the underlying infrastructure capacity and migration of customer workloads is progressing well.
We are observing substantial improvements in control plane stability and performance. The recurring disruption patterns described in earlier updates are no longer present.
Migration work will continue throughout the day. We will provide an update once migrations are completed or if any changes in status occur.
Aug 6, 08:04 UTC
Update -
Memory limit adjustments and migration have had positive effects on the first control plane cluster. Customer situated in the first control plane cluster should already see substantial improvements in performance and stability.
We have started to roll out memory adjustments in the remaining control plane cluster, as well. Rolling out the memory limits can lead to temporary unavailability of affected control planes. These interruptions should be brief and will not affect running workloads.
After memory limit adjustments are fully rolled out on the second control plane cluster, horizontal scaling and migrations will resume on both control plane clusters for the next hours until workload is distributed optimally.
Aug 5, 20:43 UTC
Update -
Migration has been picked up again. We already see encouraging results after completion of first batches. In the next hours we focus on completing the migration and horizontal scaling.
During the migration, single etcds can be temporarily unavailable for time periods lasting around 30 seconds.
We expect further performance and stability improvements for all customers during and after the migration.
Aug 5, 10:07 UTC
Identified -
Migration of the first batches has been completed. The Kubernetes Team has identified a remaining issue preventing further migration. We are setting this incident back to active until the issue is resolved and the migration completed.
Aug 5, 06:23 UTC
Monitoring -
Memory adjustments have been rolled out and show positive effects.
We are starting the migrations planned to further improve control plane performance for all our customers.
We estimate that the migration will be completed in the next hours.
Control Plane performance is expected to improve already during the migration.
We are setting this incident into Monitoring status and will provide an update once the migration is completed.
Aug 4, 21:11 UTC
Update -
Our plans for further horizontal scaling of the control plane are progressing. We expect to be able to do a dry run and further testing within the next hours, before we proceed with migrations.
We aim to finish work on horizontal scaling until EOD.
We are rolling out memory configuration improvements in parallel.
Aug 4, 16:23 UTC
Update -
We are currently rolling out memory scaling measures to address recurring stability issues during compaction operations on the affected etcd clusters. Additionally, we are planning to roll out further horizontal scaling for the affected clusters today.
We will provide another update once these measures have been applied and we can assess their impact.
Aug 4, 14:37 UTC
Update -
During a service rollout today, a subset of control planes experienced temporary restarts. Affected customers may notice brief API unavailability while these control planes recover. The team is monitoring the recovery.
Separately, work on improving infrastructure capacity and load distribution continues as described in our initial update.
We will post another update once the affected control planes have fully stabilized.
Aug 4, 13:14 UTC
Identified -
We are aware of intermittent control plane unavailability affecting a subset of Managed Kubernetes customers. Affected customers may experience API call failures, deployment timeouts, and temporary disruption of cluster management operations.
Our engineering team is actively working on both immediate mitigations and longer-term architectural improvements.
Several mitigations have already been deployed, including maintenance schedule optimization, compaction regression fixes, and storage performance improvements. Additional measures - including infrastructure migration, dedicated event etcd clusters, improved load balancing, and horizontal scaling - are in progress.
We are providing regular updates on this page. Customers experiencing issues are encouraged to subscribe to this incident for timely notifications.
Aug 4, 10:24 UTC