3 papers
cs.CL2026
The Hypocrisy Gap: Quantifying Divergence Between Internal Belief and Chain-of-Thought Explanation via Sparse Autoencoders
Shikhar Shiromani, Archie Chaudhury, Sri Pranav Kunda
Large Language Models (LLMs) frequently exhibit unfaithful behavior, producing a final answer that differs significantly from their internal chain of thought (CoT) reasoning in ord…
cs.LG2025
Alignment is Localized: A Causal Probe into Preference Layers
Archie Chaudhury
Reinforcement Learning frameworks, particularly those utilizing human annotations, have become an increasingly popular method for preference fine-tuning, where the outputs of a lan…
cs.CL2025
Beyond Next Word Prediction: Developing Comprehensive Evaluation Frameworks for measuring LLM performance on real world applications
Vishakha Agrawal, Archie Chaudhury, Shreya Agrawal
While Large Language Models (LLMs) are fundamentally next-token prediction systems, their practical applications extend far beyond this basic function. From natural language proces…