1 citations · 1 across the 11 of their papers we have counts for
5 papers · 1 filter
CL-HOI: Cross-Level Human-Object Interaction Distillation from Vision Large Language Models
Jianjun Gao, Chen Cai, Ruoyu Wang +4
Human-object interaction (HOI) detection has seen advancements with Vision Language Models (VLMs), but these methods often depend on extensive manual annotations. Vision Large Lang…
Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting
Chen Cai, Zheng Wang, Jianjun Gao +4
In recent years, the rapid increase in online video content has underscored the limitations of static Video Question Answering (VideoQA) models trained on fixed datasets, as they s…
MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action Recognition
Ruoyu Wang, Wenqian Wang, Jianjun Gao +3
Driver action recognition, aiming to accurately identify drivers' behaviours, is crucial for enhancing driver-vehicle interactions and ensuring driving safety. Unlike general actio…
CM2-Net: Continual Cross-Modal Mapping Network for Driver Action Recognition
Ruoyu Wang, Chen Cai, Wenqian Wang +4
Driver action recognition has significantly advanced in enhancing driver-vehicle interactions and ensuring driving safety by integrating multiple modalities, such as infrared and d…
Video sentence grounding with temporally global textual knowledge
Cai Chen, Runzhong Zhang, Jianjun Gao +3
Temporal sentence grounding involves the retrieval of a video moment with a natural language query. Many existing works directly incorporate the given video and temporally localize…