1 paper
Junyong Choi, Cheolhyeon Park, Jaehoon Cho
Vision Transformers demonstrate remarkable global modeling capacity but often underperform in data-scarce regimes. Distilling convolutional inductive biases from a CNN teacher prov…