8 citations · 8 across the 2 of their papers we have counts for
3 papers
cs.CV2026
COMPASS: Complete Multimodal Fusion via Proxy Tokens and Shared Spaces for Ubiquitous Sensing
Hao Wang, Yanyu Qian, Pengcheng Weng +4
Missing modalities in multimodal sensing cause not only information loss but also a fusion-interface mismatch: a fusion head trained on a canonical set of modality slots must opera…
cs.CV2024
Beyond MOT: Semantic Multi-Object Tracking
Yunhao Li, Qin Li, Hao Wang +5
Current multi-object tracking (MOT) aims to predict trajectories of targets (i.e., ''where'') in videos. Yet, knowing merely ''where'' is insufficient in many crucial applications.…
cs.CV2023★ 8 cited
Collaborative Three-Stream Transformers for Video Captioning
Hao Wang, Libo Zhang, Heng Fan +1
As the most critical components in a sentence, subject, predicate and object require special attention in the video captioning task. To implement this idea, we design a novel frame…