1 citations · 1 across the 5 of their papers we have counts for
4 papers · 1 filter
Models Know Models Best: Evaluation via Model-Preferred Formats
Joonhak Lee, Sungmok Jung, Jongyeon Park +1
Performance of Large Language Models (LLMs) on multiple-choice tasks differs markedly between symbol-based and cloze-style evaluation formats. The observed discrepancies are system…
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
Sungmok Jung, Yeonkyoung So, Joonhak Lee +3
Although negation is known to challenge large language models (LLMs), benchmarks for evaluating negation understanding-especially in Korean-are scarce. We conduct a corpus-based an…
Less Is More: Reducing Token Counts Without Compromising Performance
Gyeongje Cho, Yeonkyoung So, Sangmin Lee +1
Tokenization directly affects the inference efficiency of large language models, since fragmented tokenization increases sequence length and generation cost. Although longer, multi…
Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding
Yeonkyoung So, Gyuseong Lee, Sungmok Jung +4
Negation is a fundamental linguistic phenomenon that poses ongoing challenges for Large Language Models (LLMs), particularly in tasks requiring deep semantic understanding. Current…