Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
STRIVE: Probing Reasoning Limits in Graded Plausibility Generation and Evaluation
Bhiman Kumar Baghel, Anna Chrabaszcz, Tessa Warren +3
Event knowledge concerns who does what to whom. Psycholinguists use event-plausibility judgments to examine how this knowledge supports human language processing. To isolate plausi…
cs.CL2024
Every Answer Matters: Evaluating Commonsense with Probabilistic Measures
Qi Cheng, Michael Boratko, Pranay Kumar Yelugam +4
Large language models have demonstrated impressive performance on commonsense tasks; however, these tasks are often posed as multiple-choice questions, allowing models to exploit s…