4 papers
Measure Before You Manage: Evaluating Agent Working Memory in Coding Agents
Le Chen, Zishen Wan, Baixi Sun +6
Agent working memory is heterogeneous. Objects such as instructions, artifacts, tool outputs, and agent-generated state play different semantic roles and exhibit different size, re…
CEDAR: Automata as Verifiable Interfaces for Language-Guided Embodied Action
Lekai Chen, Alvaro Velasquez, Ashutosh Trivedi
Natural-language tasking of embodied agents is rarely just goal specification: users also impose constraints that must persist while the world changes. Code-generating LLM agents c…
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning
Chih-Hsuan Yang, Jingyan Jiang, Vikram Vasudevan +7
Many math- and science-oriented agent systems use hierarchical designs with specialized reviewer roles, assuming that a dedicated review stage should help turn wrong candidates int…
Evaluating the Safety and Skill Reasoning of Large Reasoning Models Under Compute Constraints
Adarsha Balaji, Le Chen, Rajeev Thakur +2
Test-time compute scaling has demonstrated the ability to improve the performance of reasoning language models by generating longer chain-of-thought (CoT) sequences. However, this…