7 papers
Leveraging Extragradient for Effective Sharpness-Aware Minimization in Deep Learning
Yao Fu, Chunxia Zhang, Junmin Liu +3
Generalization remains a pivotal challenge in deep learning, where traditional optimizers like Stochastic Gradient Descent (SGD) often converge to sharp minima, leading to overfitt…
Zero-Order Sharpness-Aware Minimization
Yao Fu, Yihang Jin, Chunxia Zhang +3
Prompt learning has become a key method for adapting large language models to specific tasks with limited data. However, traditional gradient-based optimization methods for tuning…
Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
Quentin Anthony, Yury Tokpanov, Skyler Szot +18
We report on the first large-scale mixture-of-experts (MoE) pretraining study on pure AMD hardware, utilizing both MI300X GPUs and Pollara networking. We distill practical guidance…
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
Yinsicheng Jiang, Yao Fu, Yeqi Huang +13
The sparse Mixture-of-Experts (MoE) architecture is increasingly favored for scaling Large Language Models (LLMs) efficiently, but it depends on heterogeneous compute and memory re…
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
Yinsicheng Jiang, Yao Fu, Yeqi Huang +13
The sparse Mixture-of-Experts (MoE) architecture is increasingly favored for scaling Large Language Models (LLMs) efficiently, but it depends on heterogeneous compute and memory re…
HybridServe: Efficient Serving of Large AI Models with Confidence-Based Cascade Routing
Leyang Xue, Yao Fu, Luo Mai +1
Giant Deep Neural Networks (DNNs), have become indispensable for accurate and robust support of large-scale cloud based AI services. However, serving giant DNNs is prohibitively ex…