From the 1 of 58 linked papers with an AI index.
3 citations · 10 across the 22 of their papers we have counts for
6 papers · 1 filter
Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection
Chaoran Chen, Vy Nguyen, Ziji Zhang +7
Tool-using LLM agents are commonly trained and evaluated in environments where tool calls succeed reliably, yet deployed tools can fail transiently, persistently, or silently. Robu…
Humans' ALMANAC: A Human Collaboration Dataset of Action-Level Mental Model Annotations for Agent Collaboration
Jiaju Chen, Yuxuan Lu, Jiayi Su +10
Recent advances in LLM agents have enabled complex cognitive capabilities, such as multi-step reasoning, planning, and tool use, that increasingly position these agents as human co…
Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
Minhua Lin, Juncheng Wu, Zijun Wang +14
LLM agents are increasingly deployed as systems built around editable external harnesses, including prompts, skills, memories and tools, that shape task execution without changing…
StaRPO: Stability-Augmented Reinforcement Policy Optimization
Jinghan Zhang, Fengran Mo, Tharindu Cyril Weerasooriya +5
Reinforcement learning (RL) is effective in enhancing the accuracy of large language models in complex reasoning tasks. Existing RL policy optimization frameworks rely on final-ans…
Bridging Models to Defend: A Population-Based Strategy for Robust Adversarial Defense
Ren Wang, Yuxuan Li, Can Chen +6
Adversarial robustness is a critical measure of a neural network's ability to withstand adversarial attacks at inference time. While robust training techniques have improved defens…
SurgWound-Bench: A Benchmark for Surgical Wound Diagnosis
Jiahao Xu, Changchang Yin, Odysseas Chatzipanagiotou +6
Surgical site infection (SSI) is one of the most common and costly healthcare-associated infections and and surgical wound care remains a significant clinical challenge in preventi…