3 papers
cs.CL2025
Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
Yeongbin Seo, Dongha Lee, Jaehyung Kim +1
Autoregressive (AR) language models generate text one token at a time, which limits their inference speed. Diffusion-based language models offer a promising alternative, as they ca…
cs.CL2025
Prior-based Noisy Text Data Filtering: Fast and Strong Alternative For Perplexity
Yeongbin Seo, Gayoung Kim, Jaehyung Kim +1
As large language models (LLMs) are pretrained on massive web corpora, careful selection of data becomes essential to ensure effective and efficient learning. While perplexity (PPL…
cs.CL2025
Quantifying Genuine Awareness in Hallucination Prediction Beyond Question-Side Shortcuts
Yeongbin Seo, Dongha Lee, Jinyoung Yeo
Many works have proposed methodologies for language model (LM) hallucination detection and reported seemingly strong performance. However, we argue that the reported performance to…