4 papers
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
Jiangfei Duan, Shuo Zhang, Zerui Wang +13
Large Language Models (LLMs) like GPT and LLaMA are revolutionizing the AI industry with their sophisticated capabilities. Training these models requires vast GPU clusters and sign…
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
Huaiyuan Ying, Shuo Zhang, Linyang Li +19
The math abilities of large language models can represent their abstract reasoning ability. In this paper, we introduce and open-source our math reasoning LLMs InternLM-Math which…
CoLLiE: Collaborative Training of Large Language Models in an Efficient Way
Kai Lv, Shuo Zhang, Tianle Gu +11
Large language models (LLMs) are increasingly pivotal in a wide range of natural language processing tasks. Access to pre-trained models, courtesy of the open-source community, has…
Scaling Laws of RoPE-based Extrapolation
Xiaoran Liu, Hang Yan, Shuo Zhang +3
The extrapolation capability of Large Language Models (LLMs) based on Rotary Position Embedding is currently a topic of considerable interest. The mainstream approach to addressing…