16 citations · 20 across the 3 of their papers we have counts for
3 papers
cs.CV2022★ 1 cited
Perceiver-VL: Efficient Vision-and-Language Modeling with Iterative Latent Attention
Zineng Tang, Jaemin Cho, Jie Lei +1
We present Perceiver-VL, a vision-and-language framework that efficiently handles high-dimensional multimodal inputs such as long videos and text. Powered by the iterative latent c…
cs.CV2022★ 16 cited
TVLT: Textless Vision-Language Transformer
Zineng Tang, Jaemin Cho, Yixin Nie +1
In this work, we present the Textless Vision-Language Transformer (TVLT), where homogeneous transformer blocks take raw visual and audio inputs for vision-and-language representati…
cs.CL2021★ 3 cited
VidLanKD: Improving Language Understanding via Video-Distilled Knowledge Transfer
Zineng Tang, Jaemin Cho, Hao Tan +1
Since visual perception can give rich information beyond text descriptions for world understanding, there has been increasing interest in leveraging visual grounding for language l…