4 papers · 1 filter
FrogNano: Training a 4B Coding Agent via Online Task Synthesis
Minseon Kim, Zhengyan Shi, Emiliano Penaloza +14
We present FrogNano, a 4B coding agent designed to tackle software engineering (SWE) tasks efficiently and effectively, even under resource-constrained environments. It is post-tra…
Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight
Christopher Z. Cui, Taylor W. Killian, Prithviraj Ammanabrolu
Reasoning in Large Language Models (LLMs) poses a challenge for oversight as many misaligned behaviors do not surface until reasoning concludes. To address this, we introduce Behav…
TALES: Text Adventure Learning Environment Suite
Christopher Zhang Cui, Xingdi Yuan, Ziang Xiao +2
Reasoning is an essential skill to enable Large Language Models (LLMs) to interact with the world. As tasks become more complex, they demand increasingly sophisticated and diverse…
Thespian: Multi-Character Text Role-Playing Game Agents
Christopher Cui, Xiangyu Peng, Mark Riedl
Text-adventure games and text role-playing games are grand challenges for reinforcement learning game playing agents. Text role-playing games are open-ended environments where an a…