4 papers
Measure Before You Manage: Evaluating Agent Working Memory in Coding Agents
Le Chen, Zishen Wan, Baixi Sun +6
Agent working memory is heterogeneous. Objects such as instructions, artifacts, tool outputs, and agent-generated state play different semantic roles and exhibit different size, re…
LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks
Chih-Hsuan Yang, Jingyan Jiang, Cheng-Hau Yang +4
Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but deployment requires deciding when extra collaboration is worth its cost. We i…
Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages
Chih-Hsuan Yang, Anjir Ahmed Chowdhury, Cheng-Hau Yang +7
Multi-agent reasoning systems often use agreement, confidence, or automated scores to decide which messages should shape a final answer. Such filtering assumes that a message likel…
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning
Chih-Hsuan Yang, Jingyan Jiang, Vikram Vasudevan +7
Many math- and science-oriented agent systems use hierarchical designs with specialized reviewer roles, assuming that a dedicated review stage should help turn wrong candidates int…