3 papers
cs.AI2026
Invocation-Level Reliability of Tool-Using Agents
Afiya Noorain, Subhranshu Mohanty, Amritesh Banerjee +1
Tool-using agents fail two ways: choosing the wrong tool, or forming wrong arguments, and an early failure of either kind can silently corrupt everything downstream. We measure a c…
cs.AI2026
Measuring Cross-Task Behavioral Consistency in Language Model Agents
Amritesh Banerjee, Pranil Raichura
Agent evaluation relies almost entirely on outcome metrics such as success rate, which capture whether an agent succeeds but not how consistently it behaves. We argue that behavior…
cs.MA2026
Spectral Dynamics of Semantic Drift in Clinical Multi-Agent Language Model Networks
Amritesh Banerjee
The integration of iterative LLMs within multi-agent diagnostic frameworks requires a rigorous quantitative reevaluation of underlying communication topologies. Frequently used arc…