128 citations · 230 across the 9 of their papers we have counts for
10 papers · 1 filter
CenterCLIP: Token Clustering for Efficient Text-Video Retrieval
Shuai Zhao, Linchao Zhu, Xiaohan Wang +1
Recently, large-scale pre-training methods like CLIP have made great progress in multi-modal research such as text-video retrieval. In CLIP, transformers are vital for modeling com…
PR-RRN: Pairwise-Regularized Residual-Recursive Networks for Non-rigid Structure-from-Motion
Haitian Zeng, Yuchao Dai, Xin Yu +2
We propose PR-RRN, a novel neural-network based method for Non-rigid Structure-from-Motion (NRSfM). PR-RRN consists of Residual-Recursive Networks (RRN) and two extra regularizatio…
Less is More: Sparse Sampling for Dense Reaction Predictions
Kezhou Lin, Xiaohan Wang, Zhedong Zheng +2
Obtaining viewer responses from videos can be useful for creators and streaming platforms to analyze the video performance and improve the future user experience. In this report, w…
Connecting Language and Vision for Natural Language-Based Vehicle Retrieval
Shuai Bai, Zhedong Zheng, Xiaohan Wang +5
Vehicle search is one basic task for the efficient traffic management in terms of the AI City. Most existing practices focus on the image-based vehicle matching, including vehicle…
T2VLAD: Global-Local Sequence Alignment for Text-Video Retrieval
Xiaohan Wang, Linchao Zhu, Yi Yang
Text-video retrieval is a challenging task that aims to search relevant video contents based on natural language descriptions. The key to this problem is to measure text-video simi…
Learning to Anticipate Egocentric Actions by Imagination
Yu Wu, Linchao Zhu, Xiaohan Wang +2
Anticipating actions before they are executed is crucial for a wide range of practical applications, including autonomous driving and robotics. In this paper, we study the egocentr…