2 papers
cs.DC2026
PRISM: Evaluating POSIX Storage Systems for AI Research Workflows
Adithya Kumar, Aditya Basu, Jacob Kahn +3
The rapid advancement of AI research is driven by massive investments in GPU clusters, yet the critical role of storage systems in enabling efficient research workflows is often ov…
cs.DC2025
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
Apostolos Kokolis, Michael Kuchnik, John Hoffman +7
Reliability is a fundamental challenge in operating large-scale machine learning (ML) infrastructures, particularly as the scale of ML models and training clusters continues to gro…