19 papers
The Hard Decision Layer: Evidence for Committed Inference in Transformers
Ashwath Vaithinathan Aravindan, Mayank Kejriwal
We investigate where and how transformer-based language models commit to predictions in multiple-choice question answering. We identify the _Hard Decision Layer_ (HDL), a natural a…
Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models
Ashwath Vaithinathan Aravindan, Mayank Kejriwal
Large language models often generate confident but incorrect answers rather than abstaining when uncertain. This problem is particularly acute for small language models (SLMs), whe…
A Compound AI Agent for Conversational Grant Discovery
Zhisheng Tang, Mayank Kejriwal
Research funding discovery remains fundamentally fragmented: researchers navigate disparate agency portals (e.g., in the United States, NSF, NIH, DARPA, Grants.gov, and many others…
Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review
Vaibhava Lakshmi Ravideshik, Mayank Kejriwal
AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and comparing the quality of AI-generat…
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
Ashwath Vaithinathan Aravindan, Mayank Kejriwal
Chain-of-Thought (CoT) prompting has emerged as a foundational technique for eliciting reasoning from Large Language Models (LLMs), yet the robustness of this approach to corruptio…
ClinicBot: A Guideline-Grounded Clinical Chatbot with Prioritized Evidence RAG and Verifiable Citations
Navapat Nananukul, Mayank Kejriwal
Clinical diagnosis requires answers that are accurate, verifiable, and explicitly grounded in official guidelines. While large language models excel at natural language processing,…