collaborators

8 papers

cs.AI2026

The Hard Decision Layer: Evidence for Committed Inference in Transformers

Ashwath Vaithinathan Aravindan, Mayank Kejriwal

We investigate where and how transformer-based language models commit to predictions in multiple-choice question answering. We identify the _Hard Decision Layer_ (HDL), a natural a…

cs.AI2026

Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models

Ashwath Vaithinathan Aravindan, Mayank Kejriwal

Large language models often generate confident but incorrect answers rather than abstaining when uncertain. This problem is particularly acute for small language models (SLMs), whe…

cs.CL2026

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations

Ashwath Vaithinathan Aravindan, Mayank Kejriwal

Chain-of-Thought (CoT) prompting has emerged as a foundational technique for eliciting reasoning from Large Language Models (LLMs), yet the robustness of this approach to corruptio…

cs.CL2026

Rewarding Intellectual Humility Learning When Not To Answer In Large Language Models

Abha Jha, Akanksha Mahajan, Ashwath Vaithinathan Aravindan +3

Large Language Models (LLMs) often produce hallucinated or unverifiable content, undermining their reliability in factual domains. This work investigates Reinforcement Learning wit…

cs.CV2025

Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability

Ashwath Vaithinathan Aravindan, Abha Jha, Mihir Kulkarni

Vision-Language Models (VLMs) have shown remarkable performance in integrating visual and textual information for tasks such as image captioning and visual question answering. Howe…

cs.CV2025

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation

Ashwath Vaithinathan Aravindan, Abha Jha, Matthew Salaway +2

Text-to-image diffusion models have revolutionized generative AI, but their vulnerability to backdoor attacks poses significant security risks. Adversaries can inject imperceptible…