3 papers
stat.ME2025
Handling Missing Responses under Cluster Dependence with Applications to Language Model Evaluation
Zhenghao Zeng, David Arbour, Avi Feller +3
Human annotations play a crucial role in evaluating the performance of GenAI models. Two common challenges in practice, however, are missing annotations (the response variable of i…
cs.CL2024
Machine Psychology
Thilo Hagendorff, Ishita Dasgupta, Marcel Binz +5
Large language models (LLMs) show increasingly advanced emergent capabilities and are being incorporated across various societal domains. Understanding their behavior and reasoning…
cs.CL2024
Language models show human-like content effects on reasoning tasks
Ishita Dasgupta, Andrew K. Lampinen, Stephanie C. Y. Chan +5
Reasoning is a key ability for an intelligent system. Large language models (LMs) achieve above-chance performance on abstract reasoning tasks, but exhibit many imperfections. Howe…