3 papers
cs.DC2025
Themis: Efficient Sparse Model Training Through Fully Sharded Sparse Data Parallelism
Yuhao Qing, Guichao Zhu, Fanxin Li +10
Mixture-of-Experts (MoE) scales large language models cost-effectively, but expert-parallel training suffers severe straggler effects from skewed expert loads. Current systems freq…
cs.CL2024
XL3M: A Training-free Framework for LLM Length Extension Based on Segment-wise Inference
Shengnan Wang, Youhui Bai, Lin Zhang +7
Length generalization failure problem, namely the large language model (LLM) fails to generalize to texts longer than its maximum training length, greatly restricts the application…
cs.RO2024
AGRNav: Efficient and Energy-Saving Autonomous Navigation for Air-Ground Robots in Occlusion-Prone Environments
Junming Wang, Zekai Sun, Xiuxian Guan +6
The exceptional mobility and long endurance of air-ground robots are raising interest in their usage to navigate complex environments (e.g., forests and large buildings). However,…