3 papers
cs.CL2026
Models Know Models Best: Evaluation via Model-Preferred Formats
Joonhak Lee, Sungmok Jung, Jongyeon Park +1
Performance of Large Language Models (LLMs) on multiple-choice tasks differs markedly between symbol-based and cloze-style evaluation formats. The observed discrepancies are system…
cs.CL2025
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
Chanwoo Park, Suyoung Park, JiA Kang +6
We present Ko-MuSR, the first benchmark to comprehensively evaluate multistep, soft reasoning in long Korean narratives while minimizing data contamination. Built following MuSR, K…
cs.CL2025
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
Chanwoo Park, Suyoung Park, Yelim Ahn +3
While traditional line-level filtering techniques, such as line-level deduplication and trailing-punctuation filters, are commonly used, these basic methods can sometimes discard v…