2 papers
cs.CL2026
Benchmarks Are Not That Out of Distribution: Word Overlap Predicts Performance
Woojin Chung, Jeonghoon Kim
Understanding what constitutes high-quality pre-training data remains a central question in language model training. In this work, we investigate whether benchmark performance is p…
cs.CL2025
Exploiting Vocabulary Frequency Imbalance in Language Model Pre-training
Woojin Chung, Jeonghoon Kim
Large language models are trained with tokenizers, and the resulting token distribution is highly imbalanced: a few words dominate the stream while most occur rarely. Recent practi…