20 citations · 44 across the 3 of their papers we have counts for
5 papers · 1 filter
Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering
Ting Yu, Kunhao Fu, Shuhui Wang +2
Video Question Answering (VideoQA) represents a crucial intersection between video understanding and language processing, requiring both discriminative unimodal comprehension and s…
Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
Ting Yu, Kunhao Fu, Jian Zhang +2
Long-term Video Question Answering (VideoQA) is a challenging vision-and-language bridging task focusing on semantic understanding of untrimmed long-term videos and diverse free-fo…
A Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D Scenes
Ting Yu, Xiaojun Lin, Shuhui Wang +3
Three-Dimensional (3D) dense captioning is an emerging vision-language bridging task that aims to generate multiple detailed and accurate descriptions for 3D scenes. It presents si…
ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
Zhou Yu, Dejing Xu, Jun Yu +4
Recent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to…
On Exploring Undetermined Relationships for Visual Relationship Detection
Yibing Zhan, Jun Yu, Ting Yu +1
In visual relationship detection, human-notated relationships can be regarded as determinate relationships. However, there are still large amount of unlabeled data, such as object…