8 papers
The Hard Decision Layer: Evidence for Committed Inference in Transformers
Ashwath Vaithinathan Aravindan, Mayank Kejriwal
We investigate where and how transformer-based language models commit to predictions in multiple-choice question answering. We identify the _Hard Decision Layer_ (HDL), a natural a…
Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models
Ashwath Vaithinathan Aravindan, Mayank Kejriwal
Large language models often generate confident but incorrect answers rather than abstaining when uncertain. This problem is particularly acute for small language models (SLMs), whe…
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
Ashwath Vaithinathan Aravindan, Mayank Kejriwal
Chain-of-Thought (CoT) prompting has emerged as a foundational technique for eliciting reasoning from Large Language Models (LLMs), yet the robustness of this approach to corruptio…
Rewarding Intellectual Humility Learning When Not To Answer In Large Language Models
Abha Jha, Akanksha Mahajan, Ashwath Vaithinathan Aravindan +3
Large Language Models (LLMs) often produce hallucinated or unverifiable content, undermining their reliability in factual domains. This work investigates Reinforcement Learning wit…
Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability
Ashwath Vaithinathan Aravindan, Abha Jha, Mihir Kulkarni
Vision-Language Models (VLMs) have shown remarkable performance in integrating visual and textual information for tasks such as image captioning and visual question answering. Howe…
Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation
Ashwath Vaithinathan Aravindan, Abha Jha, Matthew Salaway +2
Text-to-image diffusion models have revolutionized generative AI, but their vulnerability to backdoor attacks poses significant security risks. Adversaries can inject imperceptible…