Status history
Every incident and every maintenance window on every cluster is listed here, newest first. The live status is on racksystems.cloud/status. This page is the record.
An incident is any period where a reserved GPU could not start a job for a reason on our side, or where the API or console returned errors for more than 1 minute. Maintenance is announced at least 7 days ahead.
Entries
The entries below are simulated for preview and show the format of a real entry. preview
| Date | Cluster | Type | Duration | Summary |
|---|---|---|---|---|
| 2026-09-08 | C1 | Maintenance | 2 h 10 min | Host driver update to 560.35. Announced 2026-09-01. Batch jobs preempted, 3 on demand jobs finished before the window. |
| 2026-09-02 | C1 | Incident | 14 min | gpu5 reported uncorrectable ECC errors and was removed from scheduling. 1 batch job ended with reason hardware, not billed. GPU returned after reset and 30 min health check. |
| 2026-08-27 | API | Incident | 6 min | Telemetry endpoint returned 502 after a deploy. Rolled back. Job submission was not affected. |
| 2026-08-22 | C1 | Incident | 1 h 3 min | Loss of B power feed at the facility. Rack stayed up on A feed. No jobs affected. Listed for the record. |
| 2026-08-19 | C1 | Commissioned | C1 listed after 72 h soak test at 96.2 % mean utilization. |
How an entry is written
Each entry has the start time in UTC, the cluster or service affected, the duration to the minute, what happened, what it did to running jobs, and whether the minutes were billed. We write it within 1 business day of the end of the incident. If the cause is not known yet, the entry says so and is updated when it is.
Unavailable minutes from incidents feed the availability figure in the monthly capacity report and the credits in the SLA.
Subscribing
GET /v1/status/history returns this list as JSON. An RSS feed is at /status/history.xml on the main site. Reserved teams also receive an email for every incident on their cluster.