5 citations · 6 across the 3 of their papers we have counts for
3 papers · 1 filter
An Overview of Challenges in Egocentric Text-Video Retrieval
Burak Satar, Hongyuan Zhu, Hanwang Zhang +1
Text-video retrieval contains various challenges, including biases coming from diverse sources. We highlight some of them supported by illustrations to open a discussion. Besides,…
Combined CNN Transformer Encoder for Enhanced Fine-grained Human Action Recognition
Mei Chee Leong, Haosong Zhang, Hui Li Tan +2
Fine-grained action recognition is a challenging task in computer vision. As fine-grained datasets have small inter-class variations in spatial and temporal space, fine-grained act…
RoME: Role-aware Mixture-of-Expert Transformer for Text-to-Video Retrieval
Burak Satar, Hongyuan Zhu, Hanwang Zhang +1
Seas of videos are uploaded daily with the popularity of social channels; thus, retrieving the most related video contents with user textual queries plays a more crucial role. Most…