2 papers
cs.AI2026
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
Sebastián Andrés Cajas Ordóñez, Agastya Munnangi, Aldo Marzullo +13
Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a…
cs.AI2026
Testing the Black Box: Structural Barriers to Independent Evaluation of Consumer-Facing Health LLMs
Rahul Gorijavolu, Kaushik Madapati, Pritika Vig +7
Background: Consumer-facing large language models are now a common source of health information, and they interpret and personalize responses rather than retrieve them. Whether the…