activity
20192023
most citedEntroformer: A Transformer-based Entropy Model for Learned Image Compression

14 citations · 63 across the 13 of their papers we have counts for

collaborators
Showing 2023 · cs.CVShow all

5 papers · 2 filters

cs.CV2023★ 3 cited

Making Vision Transformers Efficient from A Token Sparsification View

Shuning Chang, Pichao Wang, Ming Lin +4

The quadratic computational complexity to the number of tokens limits the practical applications of Vision Transformers (ViTs). Several works propose to prune redundant tokens to a…

cs.CV2023★ 6 cited

Maximizing Spatio-Temporal Entropy of Deep 3D CNNs for Efficient Video Recognition

Junyan Wang, Zhenhong Sun, Yichen Qian +5

3D convolution neural networks (CNNs) have been the prevailing option for video recognition. To capture the temporal information, 3D convolutions are computed along the sequences,…

cs.CV2023

Aerial Diffusion: Text Guided Ground-to-Aerial View Translation from a Single Image using Diffusion Models

Divya Kothandaraman, Tianyi Zhou, Ming Lin +1

We present a novel method, Aerial Diffusion, for generating aerial views from a single ground-view image using text guidance. Aerial Diffusion leverages a pretrained text-image dif…

cs.CV2023

DeepMAD: Mathematical Architecture Design for Deep Convolutional Neural Network

Xuan Shen, Yaohua Wang, Ming Lin +4

The rapid advances in Vision Transformer (ViT) refresh the state-of-the-art performances in various vision tasks, overshadowing the conventional CNN-based models. This ignites a fe…

cs.CV2023★ 11 cited

Learning the Relation between Similarity Loss and Clustering Loss in Self-Supervised Learning

Jidong Ge, Yuxiang Liu, Jie Gui +5

Self-supervised learning enables networks to learn discriminative features from massive data itself. Most state-of-the-art methods maximize the similarity between two augmentations…