Bryan Oliver outlines chaos engineering strategies specifically designed for large-scale GPU clusters, addressing complex topologies and hardware inefficiencies. The presentation details seven practical fault-injection techniques to expose issues related to RDMA network protocols and NUMA misalignments. These methods aim to maximize the efficiency of expensive hardware while establishing robust observability loops for AI infrastructure.
- Target RDMA network protocol failures to prevent silent data corruption in distributed AI workloads.
- Inject NUMA misalignment faults to identify performance bottlenecks in multi-socket GPU servers.
- Implement seven specific fault-injection strategies tailored for large-scale cluster topologies.
- Build observability loops that correlate injected chaos with hardware efficiency metrics.
- Use chaos engineering to validate robustness of multi-million dollar GPU infrastructure.