Security of AI Training Interconnects
Threat modeling the high-speed networks that move data between accelerators in distributed AI training clusters.
Large-scale model training depends on a high-throughput fabric connecting accelerators, hosts, and storage. That performance-sensitive layer also creates security questions that conventional cloud controls do not always address directly.
This research is a collaboration with Matt Schultz and Ilil Blum Shem-Tov.
Research agenda
- Memory exposure and unauthorized data movement across RDMA-capable systems
- Multi-tenant isolation at the interconnect layer
- Security boundaries around NVIDIA GPUDirect and NCCL/RCCL communication patterns
- Detection opportunities for misuse without undermining training performance
The current research agenda organizes these questions into trust boundaries, abuse cases, and potential control points. The goal is to translate network, systems, and cloud-security principles into practical controls for large-scale AI infrastructure.