18 citations · 19 across the 6 of their papers we have counts for
4 papers · 1 filter
Skrull: Towards Efficient Long Context Fine-tuning through Dynamic Data Scheduling
Hongtao Xu, Wenting Shen, Yuanxin Wei +6
Long-context supervised fine-tuning (Long-SFT) plays a vital role in enhancing the performance of large language models (LLMs) on long-context tasks. To smoothly adapt LLMs to long…
ROAM: memory-efficient large DNN training via optimized operator ordering and memory layout
Huiyao Shu, Ang Wang, Ziji Shi +3
As deep learning models continue to increase in size, the memory requirements for training have surged. While high-level techniques like offloading, recomputation, and compression…
M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining
Junyang Lin, An Yang, Jinze Bai +9
Recent expeditious developments in deep learning algorithms, distributed training, and even hardware design for large models have enabled training extreme-scale models, say GPT-3 a…
M6-T: Exploring Sparse Expert Models and Beyond
An Yang, Junyang Lin, Rui Men +12
Mixture-of-Experts (MoE) models can achieve promising results with outrageous large amount of parameters but constant computation cost, and thus it has become a trend in model scal…