2 papers
cs.CL2026
SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue
Stephanie Fong, Yiwen Jiang, Zimu Wang +12
Large Language Models (LLMs) are increasingly used in advice seeking and decision making that may affect social judgements. Despite stigma's profound effects on people and communit…
cs.AI2026
VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don't Mean Preferences
Yiwen Jiang, Yang Deng, Stephanie Fong +9
Personalized Large Language Models (PLLMs) aim to tailor responses to individual users, where a central challenge is preference reasoning: inferring query-relevant preferences from…