5 papers
Direct Model State Migration for Elastic Training of Large Language Models
Weijian Liu, Mingzhen Li, Rui Kang +3
Large language model (LLM) training shall adapt to dynamic resources in shared clusters to tackle the elasticity, including passive preemption and optimistic scaling. State migrati…
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
Zedong Liu, Xinyang Ma, Dejun Luo +9
LLMs are widely adopted in production, pushing inference systems to their limits. Disaggregated LLM serving (e.g., PD separation and KV state disaggregation) improves scalability a…
Compact SO(3) Equivariant Atomistic Foundation Models via Structural Pruning
Chen Wang, Siyu Hu, Guangming Tan +1
SO(3) equivariant graph neural networks have become the dominant paradigm for atomistic foundation models, achieving high accuracy and data efficiency by building rotational symmet…
A Fully GPU-Accelerated Framework for High-Performance Configuration Interaction Selection with Neural Network Quantum States
Daran Sun, Bowen Kan, Haoquan Long +13
AI-driven methods have demonstrated considerable success in tackling the central challenge of accurately solving the Schrödinger equation for complex many-body systems. Among neur…
Breaking the Training Barrier of Billion-Parameter Universal Machine Learning Interatomic Potentials
Yuanchang Zhou, Hongyu Wang, Yiming Du +12
Universal Machine Learning Interatomic Potentials (uMLIPs), pre-trained on massively diverse datasets encompassing inorganic materials and organic molecules across the entire perio…