4 papers
Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention
Yang Yu, Zhuangzhuang Chen, Lanqing Li +1
Recently, reinforcement learning (RL) has become a common choice in enhancing the reasoning capabilities of vision-language models (VLMs). Considering existing RL-based finetuning…
Scaling and Transferability of Annealing Strategies in Large Language Model Training
Siqi Wang, Zhengyu Chen, Teng Xiao +5
Learning rate scheduling is crucial for training large language models, yet understanding the optimal annealing strategies across different model configurations remains challenging…
Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs
Zhengyu Chen, Siqi Wang, Teng Xiao +5
Traditional scaling laws in natural language processing suggest that increasing model size and training data enhances performance. However, recent studies reveal deviations, partic…
Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language Models
Siqi Wang, Zhengyu Chen, Bei Li +3
The scaling of large language models (LLMs) is a critical research area for the efficiency and effectiveness of model training and deployment. Our work investigates the transferabi…