All case studies

Frontier AI Lab · Production GPU infrastructure

Making a 1,056-GPU training cluster production-ready for a frontier AI lab

The problem

A 132-node, 1,056-GPU EKS cluster needed to be production-ready for LLM training. Silent failures kept surfacing at scale: dropped EFA traffic, a storage regression capping throughput, and NCCL hangs past 64 nodes.

What changed

  • All 132 nodes / 1,056 B200 GPUs ready in production
  • GDS throughput restored ~12GB/s → ~35GB/s
  • Stable multi-node NCCL all-reduce across the full fleet