activity
20192026
most citedVideo Swin Transformer

50 citations · 97 across the 13 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2024

Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and Algorithms

Miaosen Zhang, Yixuan Wei, Zhen Xing +8

Modern vision models are trained on very large noisy datasets. While these models acquire strong capabilities, they may not follow the user's intent to output the desired results i…

cs.CV2023★ 1 cited

Human Pose as Compositional Tokens

Zigang Geng, Chunyu Wang, Yixuan Wei +3

Human pose is typically represented by a coordinate vector of body joints or their heatmap embeddings. While easy for data processing, unrealistic pose estimates are admitted due t…

cs.CV2022★ 12 cited

Exploring Discrete Diffusion Models for Image Captioning

Zixin Zhu, Yixuan Wei, Jianfeng Wang +7

The image captioning task is typically realized by an auto-regressive method that decodes the text tokens one by one. We present a diffusion-based captioning model, dubbed the name…

cs.CV2022

Attentive Mask CLIP

Yifan Yang, Weiquan Huang, Yixuan Wei +8

Image token removal is an efficient augmentation strategy for reducing the cost of computing image features. However, this efficient augmentation strategy has been found to adverse…

cs.CV2022★ 7 cited

iCAR: Bridging Image Classification and Image-text Alignment for Visual Recognition

Yixuan Wei, Yue Cao, Zheng Zhang +4

Image classification, which classifies images by pre-defined categories, has been the dominant approach to visual representation learning over the last decade. Visual learning thro…

cs.CV2021★ 50 cited

Video Swin Transformer

Ze Liu, Jia Ning, Yue Cao +4

The vision community is witnessing a modeling shift from CNNs to Transformers, where pure Transformer architectures have attained top accuracy on the major video recognition benchm…