activity
20152021
most citedConvNet Architecture Search for Spatiotemporal Feature Learning

348 citations · 839 across the 25 of their papers we have counts for

collaborators

50 papers

cs.CV20213 cited

Joint Multimedia Event Extraction from Video and Article

Brian Chen, Xudong Lin, Christopher Thomas +5

Visual and textual modalities contribute complementary information about events described in multimedia documents. Videos contain rich dynamics and detailed unfoldings of events, w…

cs.CV2021

Partner-Assisted Learning for Few-Shot Image Classification

Jiawei Ma, Hanchen Xie, Guangxing Han +3

Few-shot Learning has been studied to mimic human visual capabilities and learn effective models without the need of exhaustive human annotation. Even though the idea of meta-learn…

cs.CV2021

Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos

Brian Chen, Andrew Rouditchenko, Kevin Duarte +10

Multimodal self-supervised learning is getting more and more attention as it allows not only to train large networks without human supervision but also to search and retrieve data…

cs.CV20211 cited

Co-Grounding Networks with Semantic Attention for Referring Expression Comprehension in Videos

Sijie Song, Xudong Lin, Jiaying Liu +2

In this paper, we address the problem of referring expression comprehension in videos, which is challenging due to complex expression and scene dynamics. Unlike previous methods wh…

cs.CV20218 cited

VX2TEXT: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs

Xudong Lin, Gedas Bertasius, Jue Wang +3

We present \textsc{Vx2Text}, a framework for text generation from multimodal inputs consisting of video plus text, speech, or audio. In order to leverage transformer networks, whic…

cs.CV20201 cited

Neuro-Symbolic Representations for Video Captioning: A Case for Leveraging Inductive Biases for Vision and Language

Hassan Akbari, Hamid Palangi, Jianwei Yang +6

Neuro-symbolic representations have proved effective in learning structure information in vision and language. In this paper, we propose a new model architecture for learning multi…