5 citations · 5 across the 2 of their papers we have counts for
1 paper · 1 filter
Yi Wang, Zhiwen Fan, Tianlong Chen +2
Vision Transformers (ViTs) have proven to be effective, in solving 2D image understanding tasks by training over large-scale image datasets; and meanwhile as a somehow separate tra…