3 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.CV2025
VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding
Henghao Zhao, Ge-Peng Ji, Rui Yan +2
The core challenge in video understanding lies in perceiving dynamic content changes over time. However, multimodal large language models struggle with temporal-sensitive video tas…
cs.CV2023★ 3 cited
DiffusionVMR: Diffusion Model for Joint Video Moment Retrieval and Highlight Detection
Henghao Zhao, Kevin Qinghong Lin, Rui Yan +1
Video moment retrieval and highlight detection have received attention in the current era of video content proliferation, aiming to localize moments and estimate clip relevances ba…