2 papers
cs.SE2026
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
Danlong Yuan, Wei Wu, Enhan Zhao +4
Reinforcement learning (RL) has become a key paradigm for training software engineering (SWE) agents, but existing pipelines typically rely on per-task containers for isolation. At…
cs.CL2025
HCR-Reasoner: Synergizing Large Language Models and Theory for Human-like Causal Reasoning
Yanxi Zhang, Xin Cong, Zhong Zhang +3
Genuine human-like causal reasoning is fundamental for strong artificial intelligence. Humans typically identify whether an event is part of the causal chain first, and then influe…