most citedMasked Autoencoders Are Scalable Vision Learners

201 citations · 294 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CV20231 cited

Bandwidth-efficient Inference for Neural Image Compression

Shanzhi Yin, Tongda Xu, Yongsheng Liang +4

With neural networks growing deeper and feature maps growing larger, limited communication bandwidth with external memory (or DRAM) and power constraints become a bottleneck in imp…

eess.IV2023

Conditional Perceptual Quality Preserving Image Compression

Tongda Xu, Qian Zhang, Yanghao Li +7

We propose conditional perceptual quality, an extension of the perceptual quality defined in \citet{blau2018perception}, by conditioning it on user defined information. Specificall…

cs.CV202361 cited

Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles

Chaitanya Ryali, Yuan-Ting Hu, Daniel Bolya +10

Modern hierarchical vision transformers have added several vision-specific components in the pursuit of supervised classification performance. While these components lead to effect…

cs.MM2023

Evaluating Strong Idempotence of Image Codec

Qian Zhang, Tongda Xu, Yanghao Li +1

In this paper, we first propose the concept of strong idempotent codec based on idempotent codec. The idempotence of codec refers to the stability of codec to re-compression. Simil…

cs.CV20232 cited

Efficient Semantic Segmentation by Altering Resolutions for Compressed Videos

Yubin Hu, Yuze He, Yanghao Li +4

Video semantic segmentation (VSS) is a computationally expensive task due to the per-frame prediction for videos of high frame rates. In recent work, compact models or adaptive net…

cs.CV2023

Reversible Vision Transformers

Karttikeya Mangalam, Haoqi Fan, Yanghao Li +4

We present Reversible Vision Transformers, a memory efficient architecture design for visual recognition. By decoupling the GPU memory requirement from the depth of the model, Reve…