3 papers
cs.LG2026
Locking Pretrained Weights via Deep Low-Rank Residual Distillation
Keitaro Sakamoto, Pierre Ablin, Federico Danieli +1
The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling their use across diverse hardwa…
cs.LG2026
Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
Keitaro Sakamoto, Issei Sato
The training dynamics of deep neural networks often defy expectations, even as these models form the foundation of modern machine learning. Two prominent examples are grokking, whe…
cs.LG2025
Benign Overfitting in Token Selection of Attention Mechanism
Keitaro Sakamoto, Issei Sato
Attention mechanism is a fundamental component of the transformer model and plays a significant role in its success. However, the theoretical understanding of how attention learns…