All case studies
Frontier AI Lab · Production GPU infrastructure
Making a 1,056-GPU training cluster production-ready for a frontier AI lab
The problem
A 132-node, 1,056-GPU EKS cluster needed to be production-ready for LLM training. Silent failures kept surfacing at scale: dropped EFA traffic, a storage regression capping throughput, and NCCL hangs past 64 nodes.
What changed
- All 132 nodes / 1,056 B200 GPUs ready in production
- GDS throughput restored ~12GB/s → ~35GB/s
- Stable multi-node NCCL all-reduce across the full fleet