Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
A Parallel Alternative for Energy-Efficient Neural Network Training and Inferencing
Sudip K. Seal, Maksudul Alam, Jorge Ramirez +2
Energy efficiency of training and inferencing with large neural network models is a critical challenge facing the future of sustainable large-scale machine learning workloads. This…
cs.LG2025
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
Yueming Yuan, Ahan Gupta, Jianping Li +3
Emerging expert-specialized Mixture-of-Experts (MoE) architectures, such as DeepSeek-MoE, deliver strong model quality through fine-grained expert segmentation and large top-k rout…
cs.LG2024
Scalable Artificial Intelligence for Science: Perspectives, Methods and Exemplars
Wesley Brewer, Aditya Kashi, Sajal Dash +4
In a post-ChatGPT world, this paper explores the potential of leveraging scalable artificial intelligence for scientific discovery. We propose that scaling up artificial intelligen…