69 citations · 71 across the 3 of their papers we have counts for
3 papers
cs.CV2022★ 69 cited
OmniVL:One Foundation Model for Image-Language and Video-Language Tasks
Junke Wang, Dongdong Chen, Zuxuan Wu +7
This paper presents OmniVL, a new foundation model to support both image-language and video-language tasks using one universal architecture. It adopts a unified transformer-based v…
cs.CV2022★ 2 cited
When Shift Operation Meets Vision Transformer: An Extremely Simple Alternative to Attention Mechanism
Guangting Wang, Yucheng Zhao, Chuanxin Tang +2
Attention mechanism has been widely believed as the key to success of vision transformers (ViTs), since it provides a flexible and powerful way to model spatial relationships. Howe…
cs.SD2021
Zero-Shot Text-to-Speech for Text-Based Insertion in Audio Narration
Chuanxin Tang, Chong Luo, Zhiyuan Zhao +3
Given a piece of speech and its transcript text, text-based speech editing aims to generate speech that can be seamlessly inserted into the given speech by editing the transcript.…