7 papers
The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval
Zekai Tong, Ruiyao Xu, Aryan Shrivastava +2
Existing Large Language Model (LLM) benchmarks primarily focus on syntactically correct inputs, leaving a significant gap in evaluation on imperfect text. In this work, we study ho…
Iterative Finetuning is Mostly Idempotent
Zephaniah Roe, Jack Sanderson, Dang Nguyen +5
If a model has some behavioral tendency, such as sycophancy or misalignment, and it is trained on its own outputs, will the tendency be amplified in the next generation of models?…
Modeling and Predicting Multi-Turn Answer Instability in Large Language Models
Jiahang He, Rishi Ramachandran, Neel Ramachandran +5
As large language models (LLMs) are adopted in an increasingly wide range of applications, user-model interactions have grown in both frequency and scale. Consequently, research ha…
Know Thyself? On the Incapability and Implications of AI Self-Recognition
Xiaoyan Bai, Aryan Shrivastava, Ari Holtzman +1
Self-recognition is a crucial metacognitive capability for AI systems, relevant not only for psychological analysis but also for safety, particularly in evaluative scenarios. Motiv…
Linearly Decoding Refused Knowledge in Aligned Language Models
Aryan Shrivastava, Ari Holtzman
Most commonly used language models (LMs) are instruction-tuned and aligned using a combination of fine-tuning and reinforcement learning, causing them to refuse users requests deem…
AbsenceBench: Language Models Can't Tell What's Missing
Harvey Yiyun Fu, Aryan Shrivastava, Jared Moore +3
Large language models (LLMs) are increasingly capable of processing long inputs and locating specific information within them, as evidenced by their performance on the Needle in a…