API Issue
Updates
Summary
On July 7, 2026, between 09:32 and 11:01 UTC, a configuration error caused a partial outage affecting some of our APIs and the customer Portal.
The outage started in the AT-VIE-1, AT-VIE-2, DE-FRA-1, and DE-MUC-1 zones and was later extended to all zones as we worked to understand its full scope.
During a low risk change where we wanted to deploy a new firewall system in dry-run, a mistake led to a loss of connectivity to our internal databases.
Our orchestrators lost connection to their databases creating an outage of our APIs, alongside our Portal.
Impact
During this incident the following APIs were impacted:
- Concrete AI
- IAM
- Instance Pool
- Managed Private Network
- NLB
- Portal
- SKS
Timeline (UTC)
08:35 AM Configuration change makes it to master.
09:34 AM On-call team identifies first signs of the outage through alerting.
09:38 AM Full incident response activated. Troubleshooting to identify the root cause starts.
09:46 AM API on AT-VIE-1 and DE-MUC-1 is impacted.
09:52 AM Root cause identified in a recent firewall migration.
09:55 AM Manual correction applied to relevant infrastructure.
09:55 AM Full revert of identified configuration.
10:05 AM Mitigation is considered not sufficient. Troubleshooting continues.
10:05 AM API on all the zones is impacted.
10:07 AM Portal becomes unavailable.
10:40 AM Identified unexpected stale firewall configuration still blocking the traffic.
10:41 AM Manual correction applied to relevant infrastructure.
1044 AM Services are recovering.
10:52 AM Services are fully operational.
What happened
We are in the process of replacing our legacy firewall management by a new agent that will help us improve the efficiency of our firewall system.
The change should have brought this agent in a dry-run mode in order to validate all the firewall rules before switching the management of the firewall rules.
This procedure has been previously successfully applied on most of our infrastructure as part of this migration rollout effort.
This specific deployment of the firewall agent, targeting our database proxy layer, had already been validated in our pre-production environment and was expected to be safe to roll out further. This change should have deployed the agent in dry-run, but an unforeseen mistake led to a full deployment, setting some unfinished firewall rules blocking access to the healthcheck of our database proxy layer.
It interfered with the health checks our managed EIP addresses rely on, causing those IP addresses to be withdrawn from routing.
This made the affected database proxies unreachable, which in turn caused connection failures across every service that depends on that layer.
Our monitoring first detected the issue through failing canaries at 09:32 UTC, quickly followed by reports that our orchestration layer could not reach the database.
The symptom initially looked like a database or connection-pooling problem, so early investigation focused on that perimeter before we could confirm the database itself was healthy.
As the same pattern appeared across more zones, we widened the incident to cover all of them.
Once we traced the issue back to this change, we worked to revert it.
The mitigation was not enough and the connectivity issues persisted, further investigations were conducted in order to understand the issue.
The rollback did not trigger a full cleanup of the firewall configurations and rules, explaining why the connectivity issue persisted.
As our internal tooling and automation system was impacted too, we had to manually flush these faulty rules on every database proxy instance.
By that time the orchestrators managed to regain database connectivity and the API started to be available again.
By 10:52 UTC all the services were confirmed working again, and by 11:01 UTC we considered the incident resolved across all affected zones.
What we learned
The validation we had in pre-production did not carry over cleanly to production, and we did not have a way to catch that gap before the change went out.
We also did not have a pre-tested rollback procedure ready for this specific change, so when our normal automated path was unavailable, we had to workaround with a manual fix under time pressure.
Finally, because we assessed this change as low risk due to its dry-run nature, we rolled it out more broadly than we should have. A staged, single-zone rollout would have contained the impact significantly.
What we’ve changed
We are reworking the deployment procedure to explain how to fully rollback and ensure no leftovers remain.
We are also continuing working on improving the visibility of our stack to ensure nothing is missing.
We sincerely apologize to all customers who experienced disruption during this incident. We take incidents like this seriously, given the real impact on the workloads and people relying on our platform. We’re committed to following through on the changes above and ensure it doesn’t happen again.
Recovery is completed. Services are back to nominal state. We are monitoring the situation
The services recovery is still in progress. We are monitoring the recovery and the situation
Our mitigation wasn’t effective. We continue working towards a resolution.
Some IAM operation like key creation or changes is also impacted by the API issue
The portal is also affected and currently unavailable.
We continue working towards migration
We are extending impact to additional affected zones CH-GVA2,CH-DK-2.BG-SOF-1,HR-ZAG-1
We are currently experiencing an issue with the API, which is unavailable for the following services:
- NLB
- SKS
- Managed private network
- Instance pool
The issue is being investigated
← Back