3 papers
cs.CL2026
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
Sungmok Jung, Yeonkyoung So, Joonhak Lee +3
Although negation is known to challenge large language models (LLMs), benchmarks for evaluating negation understanding-especially in Korean-are scarce. We conduct a corpus-based an…
cs.CL2026
Models Know Models Best: Evaluation via Model-Preferred Formats
Joonhak Lee, Sungmok Jung, Jongyeon Park +1
Performance of Large Language Models (LLMs) on multiple-choice tasks differs markedly between symbol-based and cloze-style evaluation formats. The observed discrepancies are system…
cs.CL2026
Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding
Yeonkyoung So, Gyuseong Lee, Sungmok Jung +4
Negation is a fundamental linguistic phenomenon that poses ongoing challenges for Large Language Models (LLMs), particularly in tasks requiring deep semantic understanding. Current…