works on

From the 1 of 33 linked papers with an AI index.

activity
20242026
collaborators

33 papers

cs.CL2026

Pigeonholing: how bad prompts hurt models, causing collapse and mistakes

Hyunji Nam, Keertana Chidambaram, Dorottya Demszky +1

The paper studies how poorly chosen prompts and conversation contexts cause large language models to repeat mistakes, narrow their output diversity, and change stances—a problem th…

cs.LG2026

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

Mickel Liu, Liwei Jiang, Yancheng Liang +4

Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders patch exposed vulnerabilities. This s…

cs.LG2026

Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data

Hyunji Nam, Haoran Li, Natasha Jaques

While post-training has successfully improved large language models (LLMs) across a variety of domains, these gains heavily rely on human-labeled data or external verifiers. Existi…

cs.CL2026

Are Language Models Sensitive to Morally Irrelevant Distractors?

Andrew Shaw, Christina Hahn, Catherine Rasgaitis +5

With the rapid uptake of large language models (LLMs) across high-stakes settings, it is becoming increasingly important to ensure that LLMs behave in ways that align with human va…

cs.LG2026

Do as the Romans Do: Learning Universal Behaviors from Heterogeneous Agents

Caleb Chang, Davin Win Kyi, Natasha Jaques +1

Humans often acquire new skills by observing others, since observed behaviors implicitly reveal how to act in an environment. However, observations drawn from a heterogeneous popul…

cs.AR2026

Embedded Arena: Iterative Optimization via Hardware Feedback

Zhihan Zhang, Alexander Le Metzger, Jiuyang Lyu +10

Embedded devices from wildlife monitoring stations to clinical wearables require local AI inference due to latency, communication, or privacy constraints. Optimizing models for het…