postmortem [fr-par-1] - Issue with Compute Instance... Jul 22, 6:01 PM
**Incident Overview**
On July 21, a block storage cluster incident affected the Instance product on FR-PAR-1 between 20:14 and 03:45 UTC. This incident caused a regional service degradation \(FR-PAR\), resulting in the following impacts:
* Kapsule operations were unavailable on a regional level between 21:30 and 01:30 UTC.
* Instances on FR-PAR-1 were non-functional.
* All managed products relying on Instance and Kapsule experienced disruptions.
#### **Root Cause Analysis**
**5 whys**
* **Why were Instance and its dependent products unavailable?**
Because disk I/O was impossible.
* **Why was disk I/O impossible?**
Because cluster doesn't accept I/O operations.
* **Why cluster doesn’t accept I/O operations?**
Due to a fault in a cluster component \(crash\).
* **Why was a cluster component faulty?**
It suffered an OOM \(Out of Memory\) kill.
* **Why did an OOM kill occur?**
The memory limit was incorrectly configured, causing the component to consume excessive memory and triggering an OS-level OOM kill.
* **Why was the memory limit incorrectly configured?**
The cluster hardware is heterogeneous; nodes have varying memory capacities. A global setting was applied without accounting for nodes with lower memory, leading to node-level OOM kills.
* **Why was the configuration error not detected?**
We lack sufficient safeguards for this specific configuration.
### Impact on Kapsule Product
* **Why did an incident in a single Availability Zone \(AZ\) impact the entire FR-PAR region for Kapsule?**The Kapsule API was proactively disabled to prevent a cascading failure \("snowball effect"\) caused by auto-healing mechanisms reacting to widespread instance unavailability.
* **Why is a multi-AZ product affected by a single API endpoint failure?**The Kapsule API architecture is currently regional and lacks the granularity to isolate or disable specific Availability Zones. Consequently, disabling the regional API was the necessary precautionary measure to protect block storage convergence.
#### **Summary of Events**
#### **Incident Timeline \(UTC\)**
| **Time \(UTC\)** | **Event Description** |
| --- | --- |
| 20:14 UTC | OOM kill on one OSD |
| 20:28 UTC | First alert on block storage team |
| 20:43 UTC | Escalation to larger teams, incident open at company level |
| 20:48 UTC | sbs-api stops processing river jobs \(no more volume update\) |
| 21:30 UTC | API Kapsule is unavailable |
| 21:43 UTC | First restart of blk-api \(internal api, not customer facing\) |
| 22:27 UTC | Second restart of blk-api, helped unstick api calls |
| 23:48 UTC | Faulty OSD removed from production, throughput restored on cluster |
| 23:50 UTC | Teams begin relaunching operations on disk |
| 00:00 UTC | System was unavailable |
| 01:30 UTC | API Kapsule is available |
| 02:08 UTC | Beginning of second block cluster global failure |
| 03:45 UTC | Block Cluster state restored successfully, end of impact |
#### **Resolution and Improvements**
**Short-term Actions**
* During the incident timelapse, we identified a configuration issue and corrected the memory limits across all cluster nodes.
**Mid-term Actions**
* Enhance cluster configuration management to prevent environment-specific mismatches \(e.g., node-level hardware heterogeneity\).
* Implement isolation mechanisms to ensure that individual component failures do not impact the overall cluster stability.
* We also notice network saturation during the recovery leading to latencies increase, we may need to rework this part.
* Optimize network performance during recovery phases to prevent latency spikes caused by saturation.
* Improve the incident response to avoid logical bias
#### **Contact**
If you have any further questions or need assistance, please contact our support team.
resolved [KUBERNETES] - [fr-par] - Product API do... Jul 22, 6:56 AM
This incident has been resolved.
resolved [fr-par-1] - Issue with Compute Instance... Jul 22, 6:55 AM
This incident has been resolved.
monitoring [fr-par-1] - Issue with Compute Instance... Jul 22, 5:53 AM
The situation has returned to normal
monitoring [KUBERNETES] - [fr-par] - Product API do... Jul 22, 5:52 AM
The situation has returned to normal
investigating [KUBERNETES] - [fr-par] - Product API do... Jul 22, 5:30 AM
We are currently investigating this issue.
investigating [fr-par-1] - Issue with Compute Instance... Jul 22, 5:30 AM
We are currently investigating this issue.
monitoring [fr-par-1] - Issue with Compute Instance... Jul 22, 3:41 AM
The Kapsule public API is available again.
monitoring [KUBERNETES] - [fr-par] - Product API do... Jul 22, 3:40 AM
The Kapsule public API is available again.
resolved [API Gateway] - [fr-par-1] APIs increase... Jul 22, 3:33 AM
This incident has been resolved.
investigating [fr-par-1] - Issue with Compute Instance... Jul 22, 2:27 AM
The situation is stable.
All products are now operational except for the Kapsule public API, which is still unavailable.
investigating [fr-par-1] - Issue with Compute Instance... Jul 22, 1:35 AM
Disk I/O operations are starting to recover, which should begin resolving the issue for the instances.
investigating [fr-par-1] - Issue with Compute Instance... Jul 22, 12:50 AM
Our team is currently working on solving the issue.
investigating [KUBERNETES] - [fr-par] - Product API do... Jul 21, 11:43 PM
Product API is down for Kubernetes clusters
investigating [fr-par-1] - Issue with Compute Instance... Jul 21, 11:39 PM
We are still investigating this critical issue with the utmost priority.
investigating [fr-par-1] - Issue with Compute Instance... Jul 21, 11:13 PM
We are continuing to investigate this issue.
investigating [fr-par-1] - Issue with Compute Instance... Jul 21, 11:03 PM
Following a crash of one node on the block storage cluster at 20h15 UTC, some VMs are stuck.
The oncall team is investigating it.
resolved [WBHE] - [fr-par] - inconsistencies in t... Jul 21, 9:50 PM
This incident has been resolved since 19:15 UTC.
investigating [WBHE] - [fr-par] - inconsistencies in t... Jul 21, 9:09 PM
Some free domains may currently display a default hosting page or resolve incorrectly due to inconsistencies in domain mappings.
Our teams are investigating and working to restore the correct configuration.
resolved [API Gateway] - [it-mil] - connection fa... Jul 21, 6:44 PM
The incident has been resolved and the situation is stable again since 16:30 UTC.
Pushes to registry and containers deployment on the it-mil region are working again.
Thank you for your patience on the matter.
identified [API Gateway] - [it-mil] - connection fa... Jul 21, 6:11 PM
Since around 12:30 PM UTC, several internal connections to the Scaleway API within the it-mil region are failing. As a result:
- Pushes to rg.it-mil.scw.eu are failing and returning an "HTTP 500" error.
- Serverless Containers deployments are failing (resources stuck in creating)
(list is not exhaustive)
Our team is actively working on a fix as we speak, we apologise for any inconvenience caused by this incident.
resolved [BLOCK-STORAGE][FR-PAR-1] - Block Storag... Jul 21, 10:28 AM
This incident has been resolved.
identified [BLOCK-STORAGE][FR-PAR-1] - Block Storag... Jul 20, 9:07 PM
A mandatory maintenance operation targeting one of the Block Storage is taking longer than planned.
Some clients may experience latencies during write operations on sbs_5k and sbs_15K volumes.
We are monitoring it and working to tune this operation to limit impact.
resolved [API] - [it-mil-1] - Network issue API Jul 20, 4:51 PM
This incident has been resolved.
investigating [API] - [it-mil-1] - Network issue API Jul 20, 4:50 PM
Following a network issue, the product APIs were unavailable for a few minutes. The issue was immediately resolved by our teams.
resolved [Kapsule] - [All region] - Error 500 whe... Jul 20, 4:07 PM
The issue has been resolved by our product team. Pools should now be created without any further issues.
identified [Kapsule] - [All region] - Error 500 whe... Jul 20, 3:50 PM
The issue has been identified and a fix is being implemented.
investigating [Kapsule] - [All region] - Error 500 whe... Jul 20, 3:31 PM
Pool creation throw error 500 when trying to create new pool.
Our team is investigating.
identified [API Gateway] - [fr-par-1] APIs increase... Jul 20, 12:13 PM
11:13 CEST to 11:48 CEST
Users may have experienced high error rates or timeouts during this 30-minute window.
Our engineering team has identified and resolved the issue.
resolved [Serverless SQL DB] - [fr-par1] - Slow c... Jul 20, 11:38 AM
This incident has been resolved.
monitoring [Serverless SQL DB] - [fr-par1] - Slow c... Jul 17, 5:36 PM
A fix has been implemented and we are monitoring the results.
investigating [Serverless SQL DB] - [fr-par1] - Slow c... Jul 17, 12:04 PM
Customers may encounter slowdowns in database cold starts, which can take up to 30 seconds.