A GPU cluster can pass every acceptance test on delivery day and still fall over in week two. I’ve consulted on enough cluster deals to stop trusting a clean handoff on its own.
The handoff tests are a snapshot. They tell you the fabric hits its bandwidth number and the GPUs show up correctly right now. They don’t tell you which GPU starts throwing uncorrectable memory errors after six hours at full power. Or which optical link flaps the first time a real training job saturates it. Or which memory module is fine for a five-minute check and dies under a week of load. Burn-in is how you find those before you’ve paid for them.
Isn’t acceptance testing enough?
No. Acceptance testing asks “does this cluster hit its numbers right now”: bisection bandwidth, collective throughput, storage speed, the health of the worst GPU in the fleet. I wrote those gates up separately in how to tell if a GPU cluster actually works.
Burn-in asks something else: “does this cluster stay healthy under days of sustained stress.”
You want both. A snapshot catches a cluster that was built wrong. A soak catches a cluster that was built with marginal parts. That second one is more common and more expensive, because it shows up after the money has moved.
Burn-in also gives you a clean performance baseline. When a job slows down in month three, you can tell whether the hardware drifted or your code did.
What a real burn-in looks like: two tiers
A burn-in worth trusting runs in two stages. The second doesn’t start until the first is clean.
Per-host first. Each node gets stressed on its own: CPU, memory, local storage, and every GPU at sustained full load for roughly a day. That’s long enough for thermal and power faults to show up. A node that fails here never makes it into the fabric test.
Cluster-wide second. Once every node is clean on its own, the whole cluster soaks together for several days: GPU interconnect, collective operations across nodes, and the network fabric under sustained load. This tier catches the faults that only show up at scale. A link that flaps under a full all-reduce. A rail that was miscabled. A switch that drops under real traffic.
Watch it the whole time, not just at the start and end. Error counters, thermal and power throttling, node liveness. A fault that appears on hour 40 of a multi-day soak is exactly what you’re running the soak to find. This two-tier, continuously watched shape is what I’d look for in any burn-in plan.
What does a soak catch that a snapshot can’t?
Burn-in is about time and heat. The faults it finds are invisible in a quick check:
- GPUs that are fine cold and degrade hot: uncorrectable memory errors, memory that has to be remapped, hardware faults that only trip after sustained load.
- Memory and CPU that pass a short test but fault under a long one.
- Marginal optics and cabling: fabric links that stay up idle but flap or drop under a real collective.
- Throttling: clocks and power that hold for five minutes but sag after an hour at full tilt, quietly capping the performance you paid for.
- Drift: a cluster that’s measurably slower at the end of the soak than at the start. Something is already on its way out.
Every one of these is cheap to make the supplier fix during burn-in. After acceptance, the fix is your problem and your cost.
Not every finding is a stop-ship
Burn-in isn’t pass-or-scrap. Split findings into two kinds.
Hard-fail criteria are things a node must pass: uncorrectable memory errors, failed GPUs, dropped fabric links.
Scored criteria are things you can miss, but only with a written disposition: fixed now, waived on the record with a reason, or scheduled with a date.
That split keeps burn-in from turning into theater, where everything gets waved through, or paralysis, where one marginal reading blocks a cluster that otherwise works. Keep the hard line hard. Write everything else down instead of arguing it out loud.
The part that protects your money
Same theme as the rest of these posts: your leverage is highest before you sign. Two things belong in the deal before you release payment.
Proof of a clean burn-in, or the time and access to run your own. If a supplier can’t show you a burn-in result and won’t give you the cluster long enough to run one, that’s your answer. A supplier who runs real burn-ins already has the artifacts. Handing them over costs nothing.
Remediation ownership tied to the timeline. During burn-in, fixing a bad GPU or a flapping link is the supplier’s job and the supplier’s cost. After you countersign acceptance, it’s yours. That one line, who fixes what and until when, is worth more than a small move on the hourly rate. Three marginal nodes you inherit will cost you far more in lost training time than you saved at the table.
Acceptance should be a document, not a handshake. A per-node inventory matched against the purchase order. A pinned firmware and driver baseline recorded as an artifact. All hard criteria passing. Every scored finding dispositioned. Your countersignature. That certificate is what you point to in month four when a node misbehaves and everyone starts asking whose problem it is.
What to demand
- Proof of a successful burn-in, or the right and the time to run one, before acceptance and before payment.
- A two-tier soak: every node clean on its own first, then the whole fabric under sustained load for several days.
- Continuous monitoring across the whole soak, not just a reading at the start and the end.
- A hard-fail versus scored split, with every scored miss written down and dispositioned, not waved through.
- A pinned firmware and driver baseline recorded per node, so you know exactly what you accepted.
- A signed acceptance certificate, with remediation on the supplier for the duration of the burn-in.
You don’t have to run the tests yourself. You have to make burn-in a term of the deal instead of a favor you hope the supplier did. A cluster that passed a five-minute check and a cluster that survived a five-day soak are not the same purchase. Pay for the second one.