Exoscale
Platform Status

API Issue

Partial outage Global Portal CH-GVA-2 API Managed Kubernetes SKS Managed Private Networks Network Load Balancer NLB CH-DK-2 API Managed Kubernetes SKS Managed Private Networks Network Load Balancer NLB DE-FRA-1 API Managed Kubernetes SKS Managed Private Networks Network Load Balancer NLB DE-MUC-1 API Managed Kubernetes SKS Managed Private Networks Network Load Balancer NLB AT-VIE-1 API Managed Kubernetes SKS Managed Private Networks Network Load Balancer NLB AT-VIE-2 API Managed Kubernetes SKS Managed Private Networks Network Load Balancer NLB BG-SOF-1 API Managed Kubernetes SKS Managed Private Networks Network Load Balancer NLB HR-ZAG-1 API Managed Kubernetes SKS Network Load Balancer NLB
2026-07-07 11:36 CEST · 1 hour, 25 minutes

Updates

Post-mortem

Summary

On July 7, 2026, between 09:32 and 11:01 UTC, a configuration error caused a partial outage affecting some of our APIs and the customer Portal.
The outage started in the AT-VIE-1, AT-VIE-2, DE-FRA-1, and DE-MUC-1 zones and was later extended to all zones as we worked to understand its full scope.

During a low risk change where we wanted to deploy a new firewall system in dry-run, a mistake led to a loss of connectivity to our internal databases.
Our orchestrators lost connection to their databases creating an outage of our APIs, alongside our Portal.

Impact

During this incident the following APIs were impacted:

  • Concrete AI
  • IAM
  • Instance Pool
  • Managed Private Network
  • NLB
  • Portal
  • SKS

Timeline (UTC)

08:35 AM Configuration change makes it to master.
09:34 AM On-call team identifies first signs of the outage through alerting.
09:38 AM Full incident response activated. Troubleshooting to identify the root cause starts.
09:46 AM API on AT-VIE-1 and DE-MUC-1 is impacted.
09:52 AM Root cause identified in a recent firewall migration.
09:55 AM Manual correction applied to relevant infrastructure.
09:55 AM Full revert of identified configuration.
10:05 AM Mitigation is considered not sufficient. Troubleshooting continues.
10:05 AM API on all the zones is impacted.
10:07 AM Portal becomes unavailable.
10:40 AM Identified unexpected stale firewall configuration still blocking the traffic.
10:41 AM Manual correction applied to relevant infrastructure.
1044 AM Services are recovering.
10:52 AM Services are fully operational.

What happened

We are in the process of replacing our legacy firewall management by a new agent that will help us improve the efficiency of our firewall system.

The change should have brought this agent in a dry-run mode in order to validate all the firewall rules before switching the management of the firewall rules.
This procedure has been previously successfully applied on most of our infrastructure as part of this migration rollout effort.
This specific deployment of the firewall agent, targeting our database proxy layer, had already been validated in our pre-production environment and was expected to be safe to roll out further. This change should have deployed the agent in dry-run, but an unforeseen mistake led to a full deployment, setting some unfinished firewall rules blocking access to the healthcheck of our database proxy layer.
It interfered with the health checks our managed EIP addresses rely on, causing those IP addresses to be withdrawn from routing.
This made the affected database proxies unreachable, which in turn caused connection failures across every service that depends on that layer.
Our monitoring first detected the issue through failing canaries at 09:32 UTC, quickly followed by reports that our orchestration layer could not reach the database.
The symptom initially looked like a database or connection-pooling problem, so early investigation focused on that perimeter before we could confirm the database itself was healthy.
As the same pattern appeared across more zones, we widened the incident to cover all of them.
Once we traced the issue back to this change, we worked to revert it.
The mitigation was not enough and the connectivity issues persisted, further investigations were conducted in order to understand the issue.
The rollback did not trigger a full cleanup of the firewall configurations and rules, explaining why the connectivity issue persisted.
As our internal tooling and automation system was impacted too, we had to manually flush these faulty rules on every database proxy instance.

By that time the orchestrators managed to regain database connectivity and the API started to be available again.

By 10:52 UTC all the services were confirmed working again, and by 11:01 UTC we considered the incident resolved across all affected zones.

What we learned

The validation we had in pre-production did not carry over cleanly to production, and we did not have a way to catch that gap before the change went out.

We also did not have a pre-tested rollback procedure ready for this specific change, so when our normal automated path was unavailable, we had to workaround with a manual fix under time pressure.

Finally, because we assessed this change as low risk due to its dry-run nature, we rolled it out more broadly than we should have. A staged, single-zone rollout would have contained the impact significantly.

What we’ve changed

We are reworking the deployment procedure to explain how to fully rollback and ensure no leftovers remain.

We are also continuing working on improving the visibility of our stack to ensure nothing is missing.

We sincerely apologize to all customers who experienced disruption during this incident. We take incidents like this seriously, given the real impact on the workloads and people relying on our platform. We’re committed to following through on the changes above and ensure it doesn’t happen again.

July 10, 2026 · 17:26 CEST
Resolved

The issue has been resolved

July 7, 2026 · 13:01 CEST
Monitoring

Recovery is completed. Services are back to nominal state. We are monitoring the situation

July 7, 2026 · 12:52 CEST
Monitoring

The services recovery is still in progress. We are monitoring the recovery and the situation

July 7, 2026 · 12:47 CEST
Update

Mitigation applied. Services are recovering

July 7, 2026 · 12:44 CEST
Update

We continue working towards mitigation

July 7, 2026 · 12:32 CEST
Update

Our mitigation wasn’t effective. We continue working towards a resolution.

July 7, 2026 · 12:22 CEST
Investigating

Some IAM operation like key creation or changes is also impacted by the API issue

July 7, 2026 · 12:09 CEST
Investigating

The portal is also affected and currently unavailable.

We continue working towards migration

July 7, 2026 · 12:07 CEST
Update

We are extending impact to additional affected zones CH-GVA2,CH-DK-2.BG-SOF-1,HR-ZAG-1

July 7, 2026 · 12:05 CEST
Update

The issue has been identified. We are working towards mitigation.

July 7, 2026 · 11:52 CEST
Investigating

We are currently experiencing an issue with the API, which is unavailable for the following services:

  • NLB
  • SKS
  • Managed private network
  • Instance pool

The issue is being investigated

July 7, 2026 · 11:46 CEST
Issue

We are investigating an issue with the API

July 7, 2026 · 11:36 CEST

← Back