3 papers
cs.AI2026
Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools
Atul Anand, Sourav Chattaraj
Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model Context Protocol (MCP) t…
cs.SE2026
Instruction Stacking Collapse: A Benchmark and the Capability-Dependent Value of Prompt Compilation
Atul Anand, Sourav Chattaraj
Production prompts rarely carry a single instruction. One system message may require valid JSON, a word limit, three citations, and a fixed tone at the same time. We study how inst…
cs.IR2026
LLMs Remember First, Forget Last: Dual-Process Interference in Large Language Models
Sourav Chattaraj, Kanak Raj
Large language models can process millions of tokens, yet how they handle conflicting information within context remains poorly understood. From patient health logs tracking evolvi…