339 citations · 339 across the 1 of their papers we have counts for
1 paper
Hassan Akbari, Liangzhe Yuan, Rui Qian +4
We present a framework for learning multimodal representations from unlabeled data using convolution-free Transformer architectures. Specifically, our Video-Audio-Text Transformer…