1 paper · 1 filter
Aritra Mazumder, Nusrat jahan Lia
Tool-using LLM agents are mostly evaluated assuming all tools work. When a tool times out, returns a week-stale value, or has its description poisoned in deployment, the developer…