6 citations · 15 across the 11 of their papers we have counts for
3 papers · 1 filter
Provable Data Scaling Law for Meta Learning via Complexity Minimization
Kazuto Fukuchi, Ryuichiro Hataya, Kota Matsui
Pre-training has become a fundamental paradigm in modern machine learning, with one of its key empirical benefits being reduced downstream sample complexity as the scale of pre-tra…
Provable Target Sample Complexity Improvements as Pre-Trained Models Scale
Kazuto Fukuchi, Ryuichiro Hataya, Kota Matsui
Pre-trained models have become indispensable for efficiently building models across a broad spectrum of downstream tasks. The advantages of pre-trained models have been highlighted…
Self-attention Networks Localize When QK-eigenspectrum Concentrates
Han Bao, Ryuichiro Hataya, Ryo Karakida
The self-attention mechanism prevails in modern machine learning. It has an interesting functionality of adaptively selecting tokens from an input sequence by modulating the degree…