4 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.SD2022★ 4 cited
End-To-End Audiovisual Feature Fusion for Active Speaker Detection
Fiseha B. Tesema, Zheyuan Lin, Shiqiang Zhu +3
Active speaker detection plays a vital role in human-machine interaction. Recently, a few end-to-end audiovisual frameworks emerged. However, these models' inference time was not e…
cs.CV2022
TGRMPT: A Head-Shoulder Aided Multi-Person Tracker and a New Large-Scale Dataset for Tour-Guide Robot
Wen Wang, Shunda Hu, Shiqiang Zhu +5
A service robot serving safely and politely needs to track the surrounding people robustly, especially for Tour-Guide Robot (TGR). However, existing multi-object tracking (MOT) or…