50 citations · 97 across the 13 of their papers we have counts for
8 papers · 1 filter
Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and Algorithms
Miaosen Zhang, Yixuan Wei, Zhen Xing +8
Modern vision models are trained on very large noisy datasets. While these models acquire strong capabilities, they may not follow the user's intent to output the desired results i…
Human Pose as Compositional Tokens
Zigang Geng, Chunyu Wang, Yixuan Wei +3
Human pose is typically represented by a coordinate vector of body joints or their heatmap embeddings. While easy for data processing, unrealistic pose estimates are admitted due t…
Exploring Discrete Diffusion Models for Image Captioning
Zixin Zhu, Yixuan Wei, Jianfeng Wang +7
The image captioning task is typically realized by an auto-regressive method that decodes the text tokens one by one. We present a diffusion-based captioning model, dubbed the name…
Attentive Mask CLIP
Yifan Yang, Weiquan Huang, Yixuan Wei +8
Image token removal is an efficient augmentation strategy for reducing the cost of computing image features. However, this efficient augmentation strategy has been found to adverse…
iCAR: Bridging Image Classification and Image-text Alignment for Visual Recognition
Yixuan Wei, Yue Cao, Zheng Zhang +4
Image classification, which classifies images by pre-defined categories, has been the dominant approach to visual representation learning over the last decade. Visual learning thro…
Video Swin Transformer
Ze Liu, Jia Ning, Yue Cao +4
The vision community is witnessing a modeling shift from CNNs to Transformers, where pure Transformer architectures have attained top accuracy on the major video recognition benchm…