From the 1 of 33 linked papers with an AI index.
33 papers
Pigeonholing: how bad prompts hurt models, causing collapse and mistakes
Hyunji Nam, Keertana Chidambaram, Dorottya Demszky +1
The paper studies how poorly chosen prompts and conversation contexts cause large language models to repeat mistakes, narrow their output diversity, and change stances—a problem th…
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
Mickel Liu, Liwei Jiang, Yancheng Liang +4
Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders patch exposed vulnerabilities. This s…
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data
Hyunji Nam, Haoran Li, Natasha Jaques
While post-training has successfully improved large language models (LLMs) across a variety of domains, these gains heavily rely on human-labeled data or external verifiers. Existi…
Are Language Models Sensitive to Morally Irrelevant Distractors?
Andrew Shaw, Christina Hahn, Catherine Rasgaitis +5
With the rapid uptake of large language models (LLMs) across high-stakes settings, it is becoming increasingly important to ensure that LLMs behave in ways that align with human va…
Do as the Romans Do: Learning Universal Behaviors from Heterogeneous Agents
Caleb Chang, Davin Win Kyi, Natasha Jaques +1
Humans often acquire new skills by observing others, since observed behaviors implicitly reveal how to act in an environment. However, observations drawn from a heterogeneous popul…
Embedded Arena: Iterative Optimization via Hardware Feedback
Zhihan Zhang, Alexander Le Metzger, Jiuyang Lyu +10
Embedded devices from wildlife monitoring stations to clinical wearables require local AI inference due to latency, communication, or privacy constraints. Optimizing models for het…