2 papers
cs.CL2026
LLMs Can Better Capture Human Judgments--With the Right Prompts
Danica Dillion, Chen Cecilia Liu, Baihui Wang +5
Are large language models (LLMs) bad at capturing human judgment? Two commonly stated limitations are that LLMs fail to capture full distributions of responses, and that their judg…
cs.CL2024
WorldValuesBench: A Large-Scale Benchmark Dataset for Multi-Cultural Value Awareness of Language Models
Wenlong Zhao, Debanjan Mondal, Niket Tandon +3
The awareness of multi-cultural human values is critical to the ability of language models (LMs) to generate safe and personalized responses. However, this awareness of LMs has bee…