Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Polar: A Benchmark for Evaluating Political Bias in LLMs
Sangho Kim, Heejin Kim, Yoonhee Park +2
Political bias in large language models (LLMs) is increasingly significant, but difficult to measure reproducibly across political and linguistic contexts. We introduce Polar, a 4,…
cs.CL2025
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
Chanwoo Park, Suyoung Park, JiA Kang +6
We present Ko-MuSR, the first benchmark to comprehensively evaluate multistep, soft reasoning in long Korean narratives while minimizing data contamination. Built following MuSR, K…
cs.CL2025
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
Chanwoo Park, Suyoung Park, Yelim Ahn +3
While traditional line-level filtering techniques, such as line-level deduplication and trailing-punctuation filters, are commonly used, these basic methods can sometimes discard v…