context robustness 1large language models 1per-example evaluation 1prediction instability 1tail risk 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CL2026
The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context
Yanzhe Zhang, Sanmi Koyejo, Diyi Yang
The paper shows that while large language models seem robust to irrelevant context when measured by overall accuracy, adding even meaningless pseudo‑words can cause prediction flip…
cs.LG2026
Optimas: Optimizing Compound AI Systems with Globally Aligned Local Rewards
Shirley Wu, Parth Sarthi, Shiyu Zhao +10
Compound AI systems integrating multiple components, such as Large Language Models, specialized tools, and traditional machine learning models, are increasingly deployed to solve c…
cs.CL2026
HumanLM: Simulating Users with State Alignment Beats Response Imitation
Shirley Wu, Evelyn Choi, Arpandeep Khatua +7
Large Language Models (LLMs) are increasingly used to simulate how specific users respond to a given context, enabling more user-centric applications that rely on user feedback. Ho…