Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models
Yang Liu, Hongming Li, Melissa Xiaohui Qin +2
We present SemanticQA, an evaluation suite designed to assess language models (LMs) in semantic phrase processing tasks. The benchmark consolidates existing multiword expression (M…
cs.CL2026
Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty
Jingyi Ren, Ante Wang, Yunghwei Lai +5
Reliable Large Language Models (LLMs) should abstain when confidence is insufficient. However, prior studies often treat refusal as a generic "I don't know'', failing to distinguis…