5 citations · 6 across the 5 of their papers we have counts for
5 papers
Towards Debiasing Frame Length Bias in Text-Video Retrieval via Causal Intervention
Burak Satar, Hongyuan Zhu, Hanwang Zhang +1
Many studies focus on improving pretraining or developing new backbones in text-video retrieval. However, existing methods may suffer from the learning and inference bias issue, as…
Masked Diffusion with Task-awareness for Procedure Planning in Instructional Videos
Fen Fang, Yun Liu, Ali Koksal +2
A key challenge with procedure planning in instructional videos lies in how to handle a large decision space consisting of a multitude of action types that belong to various tasks.…
An Overview of Challenges in Egocentric Text-Video Retrieval
Burak Satar, Hongyuan Zhu, Hanwang Zhang +1
Text-video retrieval contains various challenges, including biases coming from diverse sources. We highlight some of them supported by illustrations to open a discussion. Besides,…
Combined CNN Transformer Encoder for Enhanced Fine-grained Human Action Recognition
Mei Chee Leong, Haosong Zhang, Hui Li Tan +2
Fine-grained action recognition is a challenging task in computer vision. As fine-grained datasets have small inter-class variations in spatial and temporal space, fine-grained act…
RoME: Role-aware Mixture-of-Expert Transformer for Text-to-Video Retrieval
Burak Satar, Hongyuan Zhu, Hanwang Zhang +1
Seas of videos are uploaded daily with the popularity of social channels; thus, retrieving the most related video contents with user textual queries plays a more crucial role. Most…