3 papers
cs.SE2026
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
Danlong Yuan, Wei Wu, Enhan Zhao +4
Reinforcement learning (RL) has become a key paradigm for training software engineering (SWE) agents, but existing pipelines typically rely on per-task containers for isolation. At…
cs.AI2026
Shorten After You're Right: Lazy Length Penalties for Reasoning RL
Danlong Yuan, Tian Xie, Shaohan Huang +5
Large reasoning models, such as OpenAI o1 or DeepSeek R1, have demonstrated remarkable performance on reasoning tasks but often incur a long reasoning path with significant memory…
cs.CL2025
ReMamba: Equip Mamba with Effective Long-Sequence Modeling
Danlong Yuan, Jiahao Liu, Bei Li +4
While the Mamba architecture demonstrates superior inference efficiency and competitive performance on short-context natural language processing (NLP) tasks, empirical evidence sug…