3 papers
cs.LG2026
Slower Generalization, Faster Memorization: A Sweet Spot in Algorithmic Learning
Shin So, Kyelim Lee, Albert No
Critical-data-size accounts of grokking suggest a natural post-threshold intuition: once training data is sufficient to identify the underlying rule, additional data should acceler…
cs.AI2026
A Theoretical Analysis of Why Masked Diffusion Models Mitigate the Reversal Curse
Moongyu Jeon, Sangwoo Shin, BumJun Kim +2
Autoregressive language models (ARMs) suffer from the reversal curse: after learning '' is ,'' they often fail on the reverse query '' is .'' Masked diffusion language…
cs.LG2025
Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
Yoonjun Cho, Soeun Kim, Dongjae Jeon +3
Decomposing weight matrices into quantization and low-rank components () is a widely used technique for compressing large lang…