3 papers
cs.CL2026
The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context
Yanzhe Zhang, Sanmi Koyejo, Diyi Yang
As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by long, partially irrelevant con…
cs.CL2026
HumanLM: Simulating Users with State Alignment Beats Response Imitation
Shirley Wu, Evelyn Choi, Arpandeep Khatua +7
Large Language Models (LLMs) are increasingly used to simulate how specific users respond to a given context, enabling more user-centric applications that rely on user feedback. Ho…
cs.LG2025
Optimas: Optimizing Compound AI Systems with Globally Aligned Local Rewards
Shirley Wu, Parth Sarthi, Shiyu Zhao +10
Compound AI systems integrating multiple components, such as Large Language Models, specialized tools, and traditional machine learning models, are increasingly deployed to solve c…