most citedA Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D Scenes

20 citations · 44 across the 3 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2024

Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering

Ting Yu, Kunhao Fu, Shuhui Wang +2

Video Question Answering (VideoQA) represents a crucial intersection between video understanding and language processing, requiring both discriminative unimodal comprehension and s…

cs.CV2024

Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering

Ting Yu, Kunhao Fu, Jian Zhang +2

Long-term Video Question Answering (VideoQA) is a challenging vision-and-language bridging task focusing on semantic understanding of untrimmed long-term videos and diverse free-fo…

cs.CV202420 cited

A Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D Scenes

Ting Yu, Xiaojun Lin, Shuhui Wang +3

Three-Dimensional (3D) dense captioning is an emerging vision-language bridging task that aims to generate multiple detailed and accurate descriptions for 3D scenes. It presents si…

cs.CV201911 cited

ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering

Zhou Yu, Dejing Xu, Jun Yu +4

Recent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to…

cs.CV201913 cited

On Exploring Undetermined Relationships for Visual Relationship Detection

Yibing Zhan, Jun Yu, Ting Yu +1

In visual relationship detection, human-notated relationships can be regarded as determinate relationships. However, there are still large amount of unlabeled data, such as object…