8 papers
A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense
Shikhar Shiromani, Leo Richter
Chain-of-thought (CoT) monitoring is meant to catch the reward hacks that look clean in the actions and betray themselves only in the reasoning. We show that this is exactly where…
Visuals Lie, Consistency Speaks: Disentangling Spatial Attention from Reliability in Vision-Language Models
Logan Mann, Yi Xia, Ajit Saravanan +6
Multimodal Foundation Models are increasingly used as reasoning agents, making reliability, knowing when a model may hallucinate, critical. A common intuition, which we call the At…
SAGE: Agentic Framework for Interpretable and Clinically Translatable Computational Pathology Biomarker Discovery
Sahar Almahfouz Nasser, Juan Francisco Pesantez Borja, Jincheng Liu +15
Engineered image-based biomarkers offer a clinically interpretable alternative to black-box AI in computational pathology, yet their discovery remains largely intuition-driven, gui…
Where Reliability Lives in Vision-Language Models: A Mechanistic Study of Attention, Hidden States, and Causal Circuits
Logan Mann, Ajit Saravanan, Ishan Dave +4
A pervasive intuition holds that vision-language models (VLMs) are most trustworthy when their attention maps look sharp: concentrated attention on the queried region should imply…
Linear Predictability of Attention Heads in Large Language Models
Khalid Shaikh, Asmit Kumar Singh, Rebecca Christopher Dsouza +1
Large language model (LLM) inference is increasingly bottlenecked by the Key-Value (KV) cache, yet the fine-grained structure of attention-head activations remains poorly understoo…
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
Rohan Subramanian Thomas, Shikhar Shiromani, Abdullah Chaudhry +4
Prompt design significantly impacts the moral competence and safety alignment of large language models (LLMs), yet empirical comparisons remain fragmented across datasets and model…