Exoscale
Platform Status

Portal and API outage

Partial outage CH-GVA-2 API CH-DK-2 API DE-FRA-1 API DE-MUC-1 API AT-VIE-1 API AT-VIE-2 API BG-SOF-1 API HR-ZAG-1 API
2026-09-01 05:57 CEST · 6 hours, 18 minutes

Updates

Post-mortem

Post-mortem

Summary

On September 1st 2026, we experienced an outage of our control plane, affecting our portal and the APIs of all our services in all zones, caused by an outdated nameserver IP address in our internal resolver configuration. Impact began around 03:20 UTC and all systems were back to nominal at 10:15 UTC.

Impact

Between 03:20 and 10:15 UTC, with the most severe impact between 03:30 and 08:30 UTC, the APIs of all our services were affected in all zones.

This was an outage of our control plane, not of the services themselves. Our API and portal returned elevated error rates and timeouts, failing intermittently rather than consistently. Because every service is managed through the API and the portal, creating, listing, modifying and deleting resources could fail or hang for any service, in any zone.

Resources already provisioned continued to operate. Compute instances kept running and existing network traffic was not interrupted. No customer data was lost or exposed.

Timeline

All times UTC.

~03:20 First errors appear in our internal services.

03:34 Our on-call team is engaged.

03:57 Errors become visible on the API and portal.

04:01 Incident opened on the status page.

04:04 Impact confirmed across several zones.

04:04 to 06:24 Investigation. Symptoms are intermittent and differ between zones and services, and initially point at several unrelated subsystems.

06:24 Name resolution identified as the common factor.

06:52 Root cause identified.

06:53 Status page updated. Incorrect configuration identified, fix being deployed.

07:00 to 08:18 Fix deployed and resolver caches flushed across all zones.

08:05 API and portal back in service.

08:29 Majority of services recovered.

09:14 Some services still degraded, monitoring continues.

10:15 All systems back to nominal.

What Happened

Our internal DNS zones are replicated to authoritative nameservers hosted outside our own infrastructure. This is a deliberate design choice made to increase resilience. Name resolution is the foundation that every other internal system depends on, so if our zones were served only from within our own platform, any failure large enough to affect that platform would also remove our ability to resolve internal names, and with it much of our ability to diagnose and recover from the failure. Replicating the zones to nameservers operated independently of our infrastructure means internal name resolution continues to work when parts of our platform do not, and keeps a path open for us to recover them.

Our resolvers therefore query both our internal nameservers and these external ones. The external nameservers are referenced by IP address in our resolver configuration.

One of those external nameserver addresses changed. The change had been announced in advance, but our resolver configuration was not updated everywhere the address appears.

That old address had since been reassigned to an unrelated network, and something at it was answering DNS queries. It had no authority over our zones, and rather than refusing the query or returning NXDOMAIN, it replied with an unrelated public IP address for every internal name it was asked about.

Had the address been left unrouted, or had it returned an error, our resolvers would have moved on to the other nameservers and the stale entry would have gone unnoticed. The entry had not been updated when the address changed, and the behaviour of the reassigned address is what resulted in an outage.

Most queries continued to be answered correctly by our internal nameservers. Only a portion reached the stale address. The result was therefore not a clean failure but an intermittent one. The same lookup could succeed, return a wrong address a moment later, then succeed again. Because a wrong answer is still a valid response, nothing retried it and nothing failed over to another nameserver, and each wrong answer was then cached for the lifetime of the record.

As those wrong answers accumulated, internal services connected to the wrong address or waited on hosts that never answered. Orchestrators lost access to their databases, our public API lost access to its backends, and service discovery lost its view of the fleet, each of them partially and unpredictably rather than all at once.

This intermittency is the main reason the incident lasted as long as it did. Symptoms differed between zones and between components, services degraded and recovered on their own, and restarting a component often appeared to fix it because its next lookup happened to be answered correctly. The investigation followed those symptoms into the private network layer and then into the databases for three hours before the pattern was recognised as a name resolution problem.

Correcting the address was a one line change. Deploying it was another story. Our configuration management and job scheduling systems depend on the same internal DNS and could not reliably distribute the fix, so it was applied manually per zone over out of band access. Correcting the configuration also did not clear the wrong answers already cached. Every local resolver cache across our infrastructure had to be flushed before resolution was reliably correct again.

What We Learned

We had no mechanism to detect that a hard-coded external nameserver address no longer matched the reality. Applying such a change depends on manual follow-up across every place the address appears, with nothing to verify that it had been completed everywhere.

Our DNS monitoring verified that names resolved, not that they resolved correctly. A wrong answer produced no alert.

A partial, intermittent failure is considerably harder to diagnose than a total one. Because most lookups still succeeded, the symptoms presented as unrelated flapping across several systems rather than as a single common cause.

Our remediation path depended on the system that had failed. Most of the time between identifying the cause and restoring service was spent working around that.

What We’ve Changed

Completed

  • We fixed the nameserver addresses in all resolver configurations, in all zones.

In progress

  • We are revisiting our monitoring strategy so that this type of problem is detected quickly, and reviewing our internal dependencies so that our recovery tooling remains usable when name resolution is degraded.

Moving Forward

DNS is a shared dependency of everything we run, and we did not treat it with the verification and ownership that implies. That is the focus of the work above.

We sincerely apologize for the disruption and thank you for your trust and patience. If you have any questions or concerns, our support team is here to help.

September 4, 2026 · 16:20 CEST
Resolved

All systems back to nominal.

September 1, 2026 · 12:15 CEST
Monitoring

The team is investigating minor remaining problems and putting in place mitigations.

September 1, 2026 · 12:00 CEST
Monitoring

Some services are still flaky. We are still actively monitoring and working on stabilising the system.

September 1, 2026 · 11:14 CEST
Monitoring

The majority of the services have recovered. The team is monitoring the situation.

September 1, 2026 · 10:29 CEST
Investigating

Deployment of config changes is complete and the team is now checking individual services and applying the changes. Services are recovering progressively.

September 1, 2026 · 10:18 CEST
Investigating

The deployment is continuing and the team is monitoring the situation. Services are recovering progressively.

September 1, 2026 · 09:45 CEST
Investigating

Changes are being deployed to all affected systems, and the team is monitoring the situation for improvement.

Our next update will be at 09h45 CEST.

September 1, 2026 · 09:16 CEST
Investigating

The team has identified an incorrect configuration and a fix is being deployed. Our next update will be at 09h15 CEST.

September 1, 2026 · 08:53 CEST
Investigating

The team is still investigating, and some corrective actions have been taken. Our next update will be at 09h00 CEST.

September 1, 2026 · 08:27 CEST
Investigating

Our team is still looking at the issue.

September 1, 2026 · 07:40 CEST
Investigating

We are still investigating the issue and wil provide an update as soon as possible.

September 1, 2026 · 06:34 CEST
Update

SKS and Block Storage seems to be impacted too.

September 1, 2026 · 06:04 CEST
Issue

We are noticing a high error rate on the API, out team is checking.

September 1, 2026 · 05:57 CEST

← Back