2 papers
cs.CR2026
PVF:Understanding AI Vulnerability Against SDCs
Xun Jiao, Fred Lin, Harish D. Dixit +8
Reliability of AI systems is a fundamental concern for the successful deployment and widespread adoption of AI technologies. Unfortunately, the escalating complexity and heterogene…
cs.DC2026
Collective Communication for 100k+ GPUs
Min Si, Pavan Balaji, Yongzhou Chen +36
The increasing scale of large language models (LLMs) necessitates highly efficient collective communication frameworks, particularly as training workloads extend to hundreds of tho…