4 papers
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
Wei Liu, Jiawei Xu, Yingru Li +4
High-quality kernel is critical for scalable AI systems, and enabling LLMs to generate such code would advance AI development. However, training LLMs for this task requires suffici…
Internalizing World Models via Self-Play Finetuning for Agentic RL
Shiqi Chen, Tongyao Zhu, Zian Wang +7
Large Language Models (LLMs) as agents often struggle in out-of-distribution (OOD) scenarios. Real-world environments are complex and dynamic, governed by task-specific rules and s…
Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs
Zhengyu Chen, Siqi Wang, Teng Xiao +5
Traditional scaling laws in natural language processing suggest that increasing model size and training data enhances performance. However, recent studies reveal deviations, partic…
Does Learning Mathematical Problem-Solving Generalize to Broader Reasoning?
Ruochen Zhou, Minrui Xu, Shiqi Chen +5
There has been a growing interest in enhancing the mathematical problem-solving (MPS) capabilities of large language models. While the majority of research efforts concentrate on c…