4 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.LG2026
Activation Compression in LLMs: Theoretical Analysis and Efficient Algorithm
Wen-Da Wei, Han-Bin Fang, Yang-Di Liu +3
Training large language models (LLMs) is highly memory-intensive, as training must store not only weights and optimizer states but also intermediate activations for backpropagation…
cs.AI2024
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
Liang Zeng, Liangjun Zhong, Liang Zhao +9
In this paper, we investigate the underlying factors that potentially enhance the mathematical reasoning capabilities of large language models (LLMs). We argue that the data scalin…
cs.CL2024★ 4 cited
Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models
Tianwen Wei, Bo Zhu, Liang Zhao +13
In this technical report, we introduce the training methodologies implemented in the development of Skywork-MoE, a high-performance mixture-of-experts (MoE) large language model (L…