activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

InterDyad: Interactive Dyadic Speech-to-Video Generation by Querying Intermediate Visual Guidance

Dongwei Pan, Longwei Guo, Jiazhi Guan +7

Despite progress in speech-to-video synthesis, existing methods often struggle to capture cross-individual dependencies and provide fine-grained control over reactive behaviors in…

cs.CV2025

Global-Local Distillation Network-Based Audio-Visual Speaker Tracking with Incomplete Modalities

Yidi Li, Yihan Li, Yixin Guo +5

In speaker tracking research, integrating and complementing multi-modal data is a crucial strategy for improving the accuracy and robustness of tracking systems. However, tracking…

cs.CV2024

STNet: Deep Audio-Visual Fusion Network for Robust Speaker Tracking

Yidi Li, Hong Liu, Bing Yang

Audio-visual speaker tracking aims to determine the location of human targets in a scene using signals captured by a multi-sensor platform, whose accuracy and robustness can be imp…

cs.CV2024

PVAFN: Point-Voxel Attention Fusion Network with Multi-Pooling Enhancing for 3D Object Detection

Yidi Li, Jiahao Wen, Bin Ren +5

The integration of point and voxel representations is becoming more common in LiDAR-based 3D object detection. However, this combination often struggles with capturing semantic inf…

cs.CV2024

Feature Completion Transformer for Occluded Person Re-identification

Tao Wang, Mengyuan Liu, Hong Liu +4

Occluded person re-identification (Re-ID) is a challenging problem due to the destruction of occluders. Most existing methods focus on visible human body parts through some prior i…