4 papers
MegaFold: Efficient Training of Next-Generation 3D Attention Protein Models on Cross-Platform GPUs
Hoa La, Ahan Gupta, Alex Morehead +2
Recent advances in biomolecular modeling have been catalyzed by models such as AlphaFold3 (AF3), which introduce science-informed changes to the transformer architecture. Unlike tr…
Gated Differential Linear Attention: A Linear-Time Decoder for High-Fidelity Medical Segmentation
Hongbo Zheng, Afshin Bozorgpour, Dorit Merhof +1
Medical image segmentation requires models that preserve fine anatomical boundaries while remaining practical for clinical deployment. Transformers capture long-range dependencies…
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
Yueming Yuan, Ahan Gupta, Jianping Li +3
Emerging expert-specialized Mixture-of-Experts (MoE) architectures, such as DeepSeek-MoE, deliver strong model quality through fine-grained expert segmentation and large top-k rout…
MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators
Beichen Huang, Yueming Yuan, Zelei Shao +1
A critical approach for efficiently deploying Mixture-of-Experts (MoE) models with massive parameters is quantization. However, state-of-the-art MoE models suffer from non-negligib…