activity
20232026
most citedOccluTrack: Rethinking Awareness of Occlusion for Enhancing Multiple Pedestrian Tracking

1 citations · 1 across the 11 of their papers we have counts for

collaborators
Showing 2024Show all

5 papers · 1 filter

cs.CV2024

CL-HOI: Cross-Level Human-Object Interaction Distillation from Vision Large Language Models

Jianjun Gao, Chen Cai, Ruoyu Wang +4

Human-object interaction (HOI) detection has seen advancements with Vision Language Models (VLMs), but these methods often depend on extensive manual annotations. Vision Large Lang…

cs.CV2024

Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting

Chen Cai, Zheng Wang, Jianjun Gao +4

In recent years, the rapid increase in online video content has underscored the limitations of static Video Question Answering (VideoQA) models trained on fixed datasets, as they s…

cs.CV2024

MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action Recognition

Ruoyu Wang, Wenqian Wang, Jianjun Gao +3

Driver action recognition, aiming to accurately identify drivers' behaviours, is crucial for enhancing driver-vehicle interactions and ensuring driving safety. Unlike general actio…

cs.CV2024

CM2-Net: Continual Cross-Modal Mapping Network for Driver Action Recognition

Ruoyu Wang, Chen Cai, Wenqian Wang +4

Driver action recognition has significantly advanced in enhancing driver-vehicle interactions and ensuring driving safety by integrating multiple modalities, such as infrared and d…

cs.CV2024

Video sentence grounding with temporally global textual knowledge

Cai Chen, Runzhong Zhang, Jianjun Gao +3

Temporal sentence grounding involves the retrieval of a video moment with a natural language query. Many existing works directly incorporate the given video and temporally localize…