What counts as downtime in a GPU cluster contract?

The provider's dashboard says the cluster was up. Your job failed anyway. Whether you get paid for that gap depends on a definition of availability buried in the contract, measured by whoever holds the logs. Read it before the outage, not during the dispute.

The provider says the GPUs were up. You say your job failed, your team couldn’t reach the cluster, and a day of a multi-week training run is gone. Both can be true at the same time.

“Up” and “usable” are different claims. Which one your contract measures decides whether that lost day is a breach worth a credit, or a conversation nobody has to have with you. In the GPU cluster deals I’ve consulted on, this is the dispute I’d bet on, and it’s completely preventable. Every term that causes it is negotiable before you sign.

Define availability where you use it, not where it’s powered on

A GPU can be powered on, enumerate correctly, and report healthy on a dashboard, and still be useless to you. The scheduler won’t place jobs on it. The network path can’t get traffic to it. The auth layer is down.

“Powered on” is the easiest thing for a provider to measure and the least useful thing for you to be paid against. Negotiate availability measured where you actually use the cluster: reachable, schedulable, and performing to a stated floor. Not just drawing power and showing green.

This becomes the whole argument once you add scale. A fleet-wide availability number blends every GPU the provider has into one percentage. A local failure (one node, one interconnect segment, one storage volume) on the exact slice you were scheduled on can be total for you and nearly invisible in the fleet number.

The bigger the pool, the smaller your outage looks in the number that decides whether you get paid.

The arithmetic: fleet-wide versus point-of-use

Take a 512-GPU cluster at an illustrative reserved rate of $2.50 per GPU-hour, over a 30-day, 720-hour month. Monthly committed spend: 512 x 720 x $2.50 = $921,600.

Here’s what two common SLA tiers allow, before anyone argues about measurement:

A 99.9% monthly SLA allows 0.1% downtime: 0.001 x 720 hours = 0.72 hours, or 43.2 minutes, for the whole month. A 99.5% monthly SLA allows 0.5% downtime: 0.005 x 720 hours = 3.6 hours, or 216 minutes, for the whole month.

Downtime allowed per month at two SLA tiers (720-hour month)Minutes allowed
99.9% SLA43.2
99.5% SLA216

Now the scenario behind most disputes. A 64-GPU slice, the one you were scheduled on, is completely unreachable for a full 24 hours. The rest of the 512-GPU fleet runs fine. From where you sit, that slice was down all day: zero availability on exactly the capacity you were paying for.

Measured fleet-wide, the same event is 64 GPUs x 24 hours = 1,536 GPU-hours of downtime, against total fleet capacity of 512 x 720 = 368,640 GPU-hours for the month. That’s 1,536 / 368,640 = about 0.417 percent of the fleet’s capacity-hours. So fleet-wide availability for the month comes out to roughly 99.583%.

Run that one number against both tiers. Against a 99.9% commitment, 99.583% is a breach, short by about 0.317 percentage points. Against a 99.5% commitment, 99.583% clears the bar. No breach at all. The fleet number never dropped far enough to trip it, even though your slice was dark for a full day.

What you actually lost on that slice: 64 GPUs x 24 hours x $2.50/hour = $3,840. That’s the real number, tied to capacity you couldn’t use. Whether you see anything close to it depends entirely on which measurement method, and which SLA tier, is sitting in the contract you signed.

Why fleet-wide measurement favors the bigger side of the table

This isn’t a coincidence. It’s how averaging works.

An outage on your slice gets diluted by every other GPU in the fleet that stayed up. The larger the fleet, the smaller your outage looks. A provider with a much bigger fleet than what you’re renting can post a good fleet-wide number every month, in good faith, while individual customers have completely different experiences depending on which slice they landed on.

So ask, specifically: is availability measured against the capacity allocated to me, or against your whole fleet? That’s the difference between the $3,840 example above being compensable and it not existing as an event the contract even recognizes.

Performance floors, not just binary uptime

A node that’s reachable but running at half its rated throughput, because of a degraded interconnect link or a throttling GPU, isn’t “down” under a binary uptime definition. It’s still costing you time on every job it touches.

A contract worth signing has a performance floor: a measured bandwidth or throughput number below which the capacity counts as degraded for SLA purposes. Without it, a provider can leave a visibly limping node in rotation forever and owe you nothing, because it never crossed the line into fully down.

Who measures, maintenance windows, and claims

Three contract mechanics decide whether any of the arithmetic above ever gets paid.

Who measures, and who holds the logs. If the only uptime record is the provider’s own monitoring, you’re arguing your claim with the other side’s evidence. Ask for independent third-party monitoring, or at minimum your own right to instrument and log availability from your side of the connection, with both logs admissible under the contract.

Maintenance windows and exclusions. Scheduled maintenance is normal and usually excluded from the SLA math. The exclusion should be bounded: a stated maximum number of hours per month, advance notice of a stated number of days, and a cap on how often it happens. An open-ended maintenance exclusion is a second SLA hiding inside the first.

Claim deadlines and evidence. Cloud-style SLAs usually require you to file within a short window, often 30 days, with timestamped logs and incident references. That process was built for a self-serve platform with thousands of customers filing small, routine claims. It fits badly on a bespoke, high-value neocloud contract, where each incident is expensive and the relationship is closer to one commercial counterparty than a self-serve account. Negotiate a claims process sized to the deal: a named contact, a defined evidence standard, and a response deadline that matches what’s at stake.

Credits are not remedies

Even a clean, well-documented breach usually gets you a service credit, a dollar amount or percentage off a future invoice. A credit is not the same as being made whole.

Back to the example. An illustrative, typical credit schedule might give 10 percent of the monthly fee for a shortfall within one percentage point of the committed threshold, scaling up for bigger shortfalls. Applied to the 99.9% breach above, 10 percent of the $921,600 monthly fee is $92,160. That happens to be more than the $3,840 of real value lost on the slice.

Looks generous. It’s an accident of how these particular numbers land, not a guarantee. Shift the fleet size, the rate, or the credit tier a little and it flips the other way just as easily. A credit schedule is a negotiated number, not a law of nature. Check it in both directions instead of assuming it’s a windfall or an insult.

The remedies that actually change your exposure are the ones to negotiate alongside, or instead of, the credit schedule: a termination right after a defined number of breaches in a rolling period, a fee step-down for the rest of the term after a serious breach, or, with a new provider, an escrowed payment or milestone-based release tied to measured performance instead of the calendar. A provider confident in its own numbers has no reason to object to being paid as those numbers are demonstrated.

The checklist, in order

  1. Confirm whether availability is defined and measured at the capacity allocated to you, or across the provider’s entire fleet, and insist on the former.
  2. Add a performance floor, not just a binary uptime test, so a degraded but technically reachable node still counts.
  3. Confirm who holds the authoritative monitoring logs, and secure your own right to log and submit evidence from your side.
  4. Bound the maintenance-window exclusion with a stated maximum duration, advance notice, and frequency cap.
  5. Negotiate a claims process sized to the deal, not a self-serve cloud form, with a named contact and a defined evidence standard.
  6. Treat the credit schedule as a starting point, not the full remedy, and add termination rights, fee step-downs, or milestone-based payment release for a new counterparty.

You don’t need to read a network diagram for any of this. You need to know which number the contract measures before the outage, because that’s the number the fight is over. For diligence on the company behind the contract, including what happens if it can’t deliver at all, see how do you vet a GPU cloud provider before you wire the deposit. For the measurement work that belongs at delivery, before any SLA clock starts, see how to tell if a GPU cluster actually works.