3 papers
cs.LG2025
Scaling and Transferability of Annealing Strategies in Large Language Model Training
Siqi Wang, Zhengyu Chen, Teng Xiao +5
Learning rate scheduling is crucial for training large language models, yet understanding the optimal annealing strategies across different model configurations remains challenging…
cs.LG2024
Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language Models
Siqi Wang, Zhengyu Chen, Bei Li +3
The scaling of large language models (LLMs) is a critical research area for the efficiency and effectiveness of model training and deployment. Our work investigates the transferabi…
cs.LG2024
Let's Ask GNN: Empowering Large Language Model for Graph In-Context Learning
Zhengyu Hu, Yichuan Li, Zhengyu Chen +4
Textual Attributed Graphs (TAGs) are crucial for modeling complex real-world systems, yet leveraging large language models (LLMs) for TAGs presents unique challenges due to the gap…