What does it actually cost to run a GPU cluster?

Jason Sun
Jason Sun · Principal Consultant
Book a call

A full build-out for a 1,152-GPU cluster, priced four honest ways, showing why moving the cost of capital from 15 to 7 percent is worth more than most of the price gaps actually being negotiated.

Everyone quotes GPU clusters in dollars per GPU-hour, as if that number were a fact about the hardware. It isn’t. For a representative 1,152-GPU build, the same silicon supports a breakeven price anywhere from about $2.78 to about $6.24 per GPU-hour, a more than 2x spread, depending on three assumptions almost nobody states out loud: how long the asset is depreciated over, what it’s assumed to be worth at the end, and what it costs to borrow the money to build it. Of those three, the last one moves the number more than any price gap you’re likely to be negotiating. That’s the headline of this post, and the rest of it is the arithmetic that gets you there.

The reference build

16 racks of Nvidia GB300 NVL72: 1,152 GPUs, roughly 135 kW per rack, about 2.16 MW of IT load. This is a generic reference configuration built from published specs, not a specific site.

LineBasisAmount
16x NVL72 racks~$5.3M/rack, includes in-rack NVLink switching~$85M
Scale-out InfiniBandXDR leaf/spine plus optics, 1,152 endpoints at 800 Gbps~$4M
North-south Ethernetnon-blocking, 400 Gbps/node class~$1.2M
Storage4.6 PB usable at 4 TB/GPU, VAST/WEKA class~$2M
Management nodes6x CPU-only control plane~$0.1M
Racks, PDU, CDU, cabling~$1M
Total~$93M
xychart-beta title "Capex by line item (~$93M total)" x-axis ["GPU racks", "InfiniBand", "Ethernet", "Storage", "Mgmt nodes", "Racks/PDU/cabling"] y-axis "USD millions" 0 --> 90 bar [85, 4, 1.2, 2, 0.1, 1]

The GPU chassis line is 85 to 90 percent of the bill of materials. That means the “everything else” line, networking, storage, power distribution, is still close to $8M ($93M minus $85M), which is the part people wave away when they say the GPUs are “basically the whole cost.”

What it costs to run, per year

Roughly $9.1M/yr, broken into four lines:

  • Power: ~$1.6M. At 2.6 MW facility load (2.16 MW IT load at PUE 1.2) and $0.07/kWh, that’s 2,600 kW x 8,760 hr x $0.07 = about $1.59M.
  • Colo: ~$4.5M
  • Managed ops: ~$2.4M
  • Transit and spares: ~$0.6M

Sum: $1.6M + $4.5M + $2.4M + $0.6M = $9.1M/yr.

Billable hours. A cluster doesn’t sell all 8,760 hours a year, maintenance, failed jobs, and idle capacity eat into it. At 85 percent utilization: 1,152 GPUs x 8,760 hr/yr x 0.85 = 8,577,792, call it 8.58M GPU-hr/yr.

$9.1M / 8.58M GPU-hr = $1.06 per GPU-hour in opex alone, before a single dollar of the $93M is paid back.

Four honest breakevens

This is where the number stops being one number. Depreciation life, residual value, and cost of capital are all judgment calls, and each one changes the floor price meaningfully. You can run these same levers against your own rack count and financing terms with the GPU ROI calculator instead of taking this reference config on faith.

ScenarioCapex recovery+ opexBreakeven $/GPU-hr
5-yr life, 20% residual, unlevered$1.72$1.06~$2.78
3-yr full recovery, unlevered$3.57$1.06~$4.63
3-yr full recovery, 15% cost of capital$3.57 + ~$1.61$1.06~$6.24
5-yr, 20% residual, 7% cost of capital$1.72 + ~$0.75$1.06~$3.53
xychart-beta title "Breakeven price under four assumption sets ($/GPU-hr)" x-axis ["5yr unlevered", "3yr unlevered", "3yr at 15%", "5yr at 7%"] y-axis "USD per GPU-hour" 0 --> 7 bar [2.78, 4.63, 6.24, 3.53]

As a sanity check, the capex-recovery lines are straightforward division: $93M x 80% (after a 20% residual) / 5 years / 8.58M GPU-hr/yr comes out to about $1.74/hr, and $93M / 3 years / 8.58M GPU-hr/yr comes out to about $3.61/hr. Both land within a couple cents of the table above; the small gap is rounding carried through each step of the underlying model, not a different assumption.

Point one: the ~$2.80 number you keep hearing is quietly a five-year number

The commonly quoted breakeven near $2.80 is the first row of that table. It requires two things nobody says out loud: a five-year useful life for the hardware, and 20 percent of the original capex still being recoverable, resold or redeployed, at the end of that life.

But the contracts actually being signed right now are three-year contracts. If the asset has to pay for itself inside the term of the contract that’s paying for it, rather than over a longer life that outlasts the deal, the floor moves to the second row: above $4.60, unlevered, with no financing cost added at all.

That’s a $1.85/hr gap between “the number everyone quotes” and “the number that’s true for a three-year deal,” and it comes entirely from a depreciation assumption, not from anything about the hardware or the deal terms. Most arguments about price are actually unnamed arguments about depreciation. If a rate being pitched to you assumes a five-year life and you’re signing a three-year contract, ask which one the number in front of you actually assumes, because it’s rarely stated.

Point two: cost of capital dominates, and this is the headline

Here’s the part that should change how you spend your negotiating time.

The financing cost added on top of straight-line capex recovery, in this simplified model, is just the outstanding capex times the annual rate, divided by billable hours: $93M x rate / 8.58M GPU-hr/yr. (This treats the full capex balance as costed at the stated rate every year rather than a declining amortization schedule; a real term loan pays down principal and would show a slightly smaller gap over time, but the order of magnitude and the direction both hold.)

At 15 percent: $93M x 0.15 / 8.58M = $1.63/hr, close to the $1.61 in the table above. At 7 percent: $93M x 0.07 / 8.58M = $0.76/hr, close to the $0.75 in the table above.

The difference: $93M x (0.15 - 0.07) / 8.58M = about $0.85 to $0.87 per GPU-hour.

That’s larger than most of the price gaps that get fought over line by line in a term sheet. A vendor holding firm on $0.20/hr, or even $0.50/hr, is a smaller number than what’s sitting on the table in the financing terms behind the deal.

The highest-leverage move available to an operator, in other words, is not winning the rate argument. It’s landing a creditworthy anchor customer whose contract is strong enough to unlock cheaper financing. A lender or lessor prices risk into the rate; a long-term, investment-grade counterparty de-risks the loan more than a better-negotiated GPU price ever will. If you’re the principal in the room, the question worth asking isn’t “can you get to $X.XX,” it’s “who is on the other side of this contract, and does that change what a lender will charge you.”

Point three: reserved and on-demand serverless are different products

This is the most common mistake in these conversations: comparing a serverless, pay-by-the-second on-demand rate to a reserved, multi-year bare-metal rate as if they were the same thing.

Modal’s public pricing page, checked September 2, 2026, lists on-demand serverless rates of $7.10/hr for B300 and $3.95/hr for H100 SXM5. Lambda’s public pricing page, also checked September 2, 2026, lists on-demand H100 SXM at $3.99 to $4.29/hr depending on cluster size. Those are pay-as-you-go, no-commitment prices, and they’re the ones that show up first in a Google search.

Reserved, multi-year bare-metal contracts price differently. Compute Exchange’s public reserved-GPU pricing page, published April 10, 2026, states that “H100 reserved pricing currently runs around $1.07 to $1.70 per hour versus on-demand rates of $2.50 to $5.00 per hour,” and separately that a one-year reserved H100 SXM contract runs near $1.89/hr and a three-year reserved H100 PCIe contract near $1.84/hr, against an on-demand H100 SXM rate of about $2.99/hr quoted on the same page.

Put the on-demand serverless number next to the reserved number for the same chip and the gap is real: $3.95/hr (Modal, serverless) against roughly $1.07 to $1.89/hr (Compute Exchange, one- to three-year reserved) works out to somewhere between about 2.1x and 3.7x, depending on exactly which reserved tier you pick. That’s the “2 to 3x apart” gap people quote past each other constantly, some of it commitment type and some of it chip generation and vendor tier, all blended into one number that gets repeated without the label. Asking “what’s the GPU rate” without specifying which product, which commitment length, and which chip generation is the fastest way to compare two numbers that were never supposed to be compared.

The takeaway, in order of what to check first

  1. Ask what depreciation life and residual value are baked into the quoted breakeven, and whether that life is longer than the contract term you’re actually signing.
  2. Ask what cost of capital is embedded in the rate, and whether a stronger anchor counterparty on your side of the deal would move that rate. This is worth more than the negotiated price itself.
  3. Ask whether you’re comparing a reserved, committed rate to an on-demand, serverless rate, or actually comparing like to like.

None of this requires being an engineer. It requires asking which assumptions are hiding inside a single dollar figure, and this post is the worked version of that question so you have the arithmetic in hand before the meeting, not after.

The GPU ROI calculator is the interactive version of that same question, letting you swap in your own contract term, residual assumption, and financing rate instead of trusting ours.