Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Single LLM Debate, MoLaCE: Mixture of Latent Concept Experts Against Confirmation Bias
Hazel Kim, Philip Torr
Large language models (LLMs) are highly vulnerable to input confirmation bias. When a prompt implies a preferred answer, models often reinforce that bias rather than explore altern…
cs.CL2025
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou +39
Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstrac…