SLA
This page summarises the service level agreement for RACK SYSTEMS clusters. The signed reservation agreement takes precedence where it differs. Preview access carries no SLA.
Availability
Reserved capacity is available 99.5 % of minutes in a calendar month, measured per reservation. A minute counts as unavailable when a reserved GPU cannot start a job for reasons on our side: hardware failure, host fault, network fault, or unannounced maintenance.
Announced maintenance windows do not count against availability. They are announced at least 7 days ahead, last at most 4 hours, and happen at most twice a month.
On demand and batch have no availability commitment. Capacity is listed when it is free. A queued job is not an outage.
Credits
If availability for a reservation falls below 99.5 % in a month, the next invoice is credited 10 % of that reservation's monthly price. Below 99.0 %, 25 %. Below 95.0 %, 50 %. Credits are the only remedy and are applied automatically, without a claim.
Hardware faults
A GPU that reports uncorrectable ECC errors, falls off the bus, or fails a health check is removed from scheduling within 60 seconds. Running jobs on it are ended with status failed and the reason hardware. Reserved teams get a replacement GPU from the on demand share of the same rack if one is free. If none is free, the minutes count as unavailable.
Metered minutes for a job ended by a hardware fault are not billed.
Data
Local NVMe is wiped at job end. Shared storage is replicated within the rack's storage cluster and backed up nightly to a second facility in the same region. Retention is the reservation term plus 7 days, or 30 days after the last on demand job. We do not read your data and we do not use it for anything.
Deletion on request completes within 7 days, backups included.
Security
Jobs run in containers under the NVIDIA runtime with no host access. GPUs are not shared between jobs. Memory is cleared between jobs. API keys are hashed at rest and shown once. Everything in transit uses TLS 1.3.
Support
Email support during business hours, Pacific time, with a 4 hour first response for reserved teams and 1 business day for on demand. A phone number for hardware incidents is provided with each reservation.
Changes
This SLA changes at most once a quarter. Changes are announced 30 days ahead and do not apply to reservations already signed.