collaborators

5 papers

cs.CL2025

Modeling and Predicting Multi-Turn Answer Instability in Large Language Models

Jiahang He, Rishi Ramachandran, Neel Ramachandran +5

As large language models (LLMs) are adopted in an increasingly wide range of applications, user-model interactions have grown in both frequency and scale. Consequently, research ha…

cs.AI2025

Know Thyself? On the Incapability and Implications of AI Self-Recognition

Xiaoyan Bai, Aryan Shrivastava, Ari Holtzman +1

Self-recognition is a crucial metacognitive capability for AI systems, relevant not only for psychological analysis but also for safety, particularly in evaluative scenarios. Motiv…

cs.CL2025

Linearly Decoding Refused Knowledge in Aligned Language Models

Aryan Shrivastava, Ari Holtzman

Most commonly used language models (LMs) are instruction-tuned and aligned using a combination of fine-tuning and reinforcement learning, causing them to refuse users requests deem…

cs.CL2025

AbsenceBench: Language Models Can't Tell What's Missing

Harvey Yiyun Fu, Aryan Shrivastava, Jared Moore +3

Large language models (LLMs) are increasingly capable of processing long inputs and locating specific information within them, as evidenced by their performance on the Needle in a…

cs.CL2025

DICE: A Framework for Dimensional and Contextual Evaluation of Language Models

Aryan Shrivastava, Paula Akemi Aoyagui

Language models (LMs) are increasingly being integrated into a wide range of applications, yet the modern evaluation paradigm does not sufficiently reflect how they are actually be…