most citedCombined CNN Transformer Encoder for Enhanced Fine-grained Human Action Recognition

5 citations · 6 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV2023

Towards Debiasing Frame Length Bias in Text-Video Retrieval via Causal Intervention

Burak Satar, Hongyuan Zhu, Hanwang Zhang +1

Many studies focus on improving pretraining or developing new backbones in text-video retrieval. However, existing methods may suffer from the learning and inference bias issue, as…

cs.CV2023

Masked Diffusion with Task-awareness for Procedure Planning in Instructional Videos

Fen Fang, Yun Liu, Ali Koksal +2

A key challenge with procedure planning in instructional videos lies in how to handle a large decision space consisting of a multitude of action types that belong to various tasks.…

cs.CV2023

An Overview of Challenges in Egocentric Text-Video Retrieval

Burak Satar, Hongyuan Zhu, Hanwang Zhang +1

Text-video retrieval contains various challenges, including biases coming from diverse sources. We highlight some of them supported by illustrations to open a discussion. Besides,…

cs.CV20225 cited

Combined CNN Transformer Encoder for Enhanced Fine-grained Human Action Recognition

Mei Chee Leong, Haosong Zhang, Hui Li Tan +2

Fine-grained action recognition is a challenging task in computer vision. As fine-grained datasets have small inter-class variations in spatial and temporal space, fine-grained act…

cs.CV20221 cited

RoME: Role-aware Mixture-of-Expert Transformer for Text-to-Video Retrieval

Burak Satar, Hongyuan Zhu, Hanwang Zhang +1

Seas of videos are uploaded daily with the popularity of social channels; thus, retrieving the most related video contents with user textual queries plays a more crucial role. Most…