Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
Wei Liu, Jiawei Xu, Yingru Li +4
High-quality kernel is critical for scalable AI systems, and enabling LLMs to generate such code would advance AI development. However, training LLMs for this task requires suffici…
cs.LG2025
Why Do LLM Agents Fail in Exploring New Environments? A World-Modeling Perspective
Shiqi Chen, Tongyao Zhu, Zian Wang +8
Large Language Models (LLMs) as agents often fail to improve in new environments. We identify and characterize a failure mode we call exploration collapse: under reinforcement lear…
cs.LG2025
Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs
Zhengyu Chen, Siqi Wang, Teng Xiao +5
Traditional scaling laws in natural language processing suggest that increasing model size and training data enhances performance. However, recent studies reveal deviations, partic…