3 papers
cs.LG2025
Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
Keitaro Sakamoto, Issei Sato
The training dynamics of deep neural networks often defy expectations, even as these models form the foundation of modern machine learning. Two prominent examples are grokking, whe…
cs.LG2024
Benign Overfitting in Token Selection of Attention Mechanism
Keitaro Sakamoto, Issei Sato
Attention mechanism is a fundamental component of the transformer model and plays a significant role in its success. However, the theoretical understanding of how attention learns…
cs.LG2024
Multiplicative Logit Adjustment Approximates Neural-Collapse-Aware Decision Boundary Adjustment
Naoya Hasegawa, Issei Sato
Real-world data distributions are often highly skewed. This has spurred a growing body of research on long-tailed recognition, aimed at addressing the imbalance in training classif…