11 citations · 13 across the 4 of their papers we have counts for
4 papers
Revisiting Multimodal Representation in Contrastive Learning: From Patch and Token Embeddings to Finite Discrete Tokens
Yuxiao Chen, Jianbo Yuan, Yu Tian +5
Contrastive learning-based vision-language pre-training approaches, such as CLIP, have demonstrated great success in many vision-language tasks. These methods achieve cross-modal a…
CAT: Causal Audio Transformer for Audio Classification
Xiaoyu Liu, Hanlin Lu, Jianbo Yuan +1
The attention-based Transformers have been increasingly applied to audio classification because of their global receptive field and ability to handle long-term dependency. However,…
HiCLIP: Contrastive Language-Image Pretraining with Hierarchy-aware Attention
Shijie Geng, Jianbo Yuan, Yu Tian +2
The success of large-scale contrastive vision-language pretraining (CLIP) has benefited both visual recognition and multimodal content understanding. The concise design brings CLIP…
Efficient Attention via Control Variates
Lin Zheng, Jianbo Yuan, Chong Wang +1
Random-feature-based attention (RFA) is an efficient approximation of softmax attention with linear runtime and space complexity. However, the approximation gap between RFA and con…