From the 1 of 35 linked papers with an AI index.
7 papers · 1 filter
Pigeonholing: how bad prompts hurt models, causing collapse and mistakes
Hyunji Nam, Keertana Chidambaram, Dorottya Demszky +1
The paper studies how poorly chosen prompts and conversation contexts cause large language models to repeat mistakes, narrow their output diversity, and change stances—a problem th…
Are Language Models Sensitive to Morally Irrelevant Distractors?
Andrew Shaw, Christina Hahn, Catherine Rasgaitis +5
With the rapid uptake of large language models (LLMs) across high-stakes settings, it is becoming increasingly important to ensure that LLMs behave in ways that align with human va…
RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
Zhiyuan Zeng, Hamish Ivison, Yiping Wang +14
We introduce Reinforcement Learning (RL) with Adaptive Verifiable Environments (RLVE), an approach using verifiable environments that procedurally generate problems and provide alg…
How LLMs Distort Our Written Language
Marwa Abdulhai, Isadora White, Yanming Wan +4
Large language models (LLMs) are used by over a billion people globally, most often to assist with writing. In this work, we demonstrate that LLMs not only alter the voice and tone…
Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
Marwa Abdulhai, Ryan Cheng, Donovan Clay +3
Large Language Models (LLMs) are increasingly used to simulate human users in interactive settings such as therapy, education, and social role-play. While these simulations enable…
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
Marwa Abdulhai, Ryan Cheng, Aryansh Shrivastava +3
Large Language Models (LLMs) interact with millions of people worldwide in applications such as customer support, education and healthcare. However, their ability to produce decept…