context robustness 1large language models 1per-example evaluation 1prediction instability 1tail risk 1
From the 1 of 3 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context
Yanzhe Zhang, Sanmi Koyejo, Diyi Yang
The paper shows that while large language models seem robust to irrelevant context when measured by overall accuracy, adding even meaningless pseudo‑words can cause prediction flip…
cs.CL2026
HumanLM: Simulating Users with State Alignment Beats Response Imitation
Shirley Wu, Evelyn Choi, Arpandeep Khatua +7
Large Language Models (LLMs) are increasingly used to simulate how specific users respond to a given context, enabling more user-centric applications that rely on user feedback. Ho…