2 papers
cs.AI2026
Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools
Atul Anand, Sourav Chattaraj
Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model Context Protocol (MCP) t…
cs.SE2026
Instruction Stacking Collapse: A Benchmark and the Capability-Dependent Value of Prompt Compilation
Atul Anand, Sourav Chattaraj
Production prompts rarely carry a single instruction. One system message may require valid JSON, a word limit, three citations, and a fixed tone at the same time. We study how inst…