14 citations · 63 across the 13 of their papers we have counts for
5 papers · 2 filters
Making Vision Transformers Efficient from A Token Sparsification View
Shuning Chang, Pichao Wang, Ming Lin +4
The quadratic computational complexity to the number of tokens limits the practical applications of Vision Transformers (ViTs). Several works propose to prune redundant tokens to a…
Maximizing Spatio-Temporal Entropy of Deep 3D CNNs for Efficient Video Recognition
Junyan Wang, Zhenhong Sun, Yichen Qian +5
3D convolution neural networks (CNNs) have been the prevailing option for video recognition. To capture the temporal information, 3D convolutions are computed along the sequences,…
Aerial Diffusion: Text Guided Ground-to-Aerial View Translation from a Single Image using Diffusion Models
Divya Kothandaraman, Tianyi Zhou, Ming Lin +1
We present a novel method, Aerial Diffusion, for generating aerial views from a single ground-view image using text guidance. Aerial Diffusion leverages a pretrained text-image dif…
DeepMAD: Mathematical Architecture Design for Deep Convolutional Neural Network
Xuan Shen, Yaohua Wang, Ming Lin +4
The rapid advances in Vision Transformer (ViT) refresh the state-of-the-art performances in various vision tasks, overshadowing the conventional CNN-based models. This ignites a fe…
Learning the Relation between Similarity Loss and Clustering Loss in Self-Supervised Learning
Jidong Ge, Yuxiang Liu, Jie Gui +5
Self-supervised learning enables networks to learn discriminative features from massive data itself. Most state-of-the-art methods maximize the similarity between two augmentations…