1 citations · 1 across the 2 of their papers we have counts for
4 papers · 1 filter
Object-Aware 4D Human Motion Generation
Shurui Gui, Deep Anil Patel, Xiner Li +1
Recent advances in video diffusion models have enabled the generation of high-quality videos. However, these videos still suffer from unrealistic deformations, semantic violations,…
Group Relative Augmentation for Data Efficient Action Detection
Deep Anil Patel, Iain Melvin, Zachary Izzo +1
Adapting large Video-Language Models (VLMs) for action detection using only a few examples poses challenges like overfitting and the granularity mismatch between scene-level pre-tr…
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection
Wentao Bao, Kai Li, Yuxiao Chen +3
Action detection aims to detect (recognize and localize) human actions spatially and temporally in videos. Existing approaches focus on the closed-set setting where an action detec…
Learning to Localize Actions in Instructional Videos with LLM-Based Multi-Pathway Text-Video Alignment
Yuxiao Chen, Kai Li, Wentao Bao +4
Learning to localize temporal boundaries of procedure steps in instructional videos is challenging due to the limited availability of annotated large-scale training videos. Recent…