2 papers
cs.LG2026
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
Mengqi Li, Lei Zhao, Anthony Man-Cho So +2
Can language models improve their reasoning performance without external rewards, using only their own sampled responses for training? We show that they can. We propose Self-evolvi…
cs.LG2025
StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs
Qijun Luo, Mengqi Li, Lei Zhao +1
Training language models on long sequence data is a demanding requirement for enhancing the model's capability on complex tasks, e.g., long-chain reasoning. However, as the sequenc…