Skip to content

Status history

Every incident and every maintenance window on every cluster is listed here, newest first. The live status is on racksystems.cloud/status. This page is the record.

An incident is any period where a reserved GPU could not start a job for a reason on our side, or where the API or console returned errors for more than 1 minute. Maintenance is announced at least 7 days ahead.

Entries

The entries below are simulated for preview and show the format of a real entry. preview

DateClusterTypeDurationSummary
2026-09-08C1Maintenance2 h 10 minHost driver update to 560.35. Announced 2026-09-01. Batch jobs preempted, 3 on demand jobs finished before the window.
2026-09-02C1Incident14 mingpu5 reported uncorrectable ECC errors and was removed from scheduling. 1 batch job ended with reason hardware, not billed. GPU returned after reset and 30 min health check.
2026-08-27APIIncident6 minTelemetry endpoint returned 502 after a deploy. Rolled back. Job submission was not affected.
2026-08-22C1Incident1 h 3 minLoss of B power feed at the facility. Rack stayed up on A feed. No jobs affected. Listed for the record.
2026-08-19C1CommissionedC1 listed after 72 h soak test at 96.2 % mean utilization.

How an entry is written

Each entry has the start time in UTC, the cluster or service affected, the duration to the minute, what happened, what it did to running jobs, and whether the minutes were billed. We write it within 1 business day of the end of the incident. If the cause is not known yet, the entry says so and is updated when it is.

Unavailable minutes from incidents feed the availability figure in the monthly capacity report and the credits in the SLA.

Subscribing

GET /v1/status/history returns this list as JSON. An RSS feed is at /status/history.xml on the main site. Reserved teams also receive an email for every incident on their cluster.

RACK SYSTEMS, us-west. Console and API access is by invitation during preview.