15 papers
UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation
Tianhao Han, Haoyang Zhang, Liang Xie +5
Manually annotating accurate 3D hand poses is extremely time-consuming and labor-intensive. Existing self-supervised hand pose estimation methods leverage the discrepancy between i…
Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios
Peizheng Yan, Yu Zhao, Liang Xie +3
Large vision-language models perform well on short- and medium-length video understanding but still struggle to maintain coherent event memory and recover long-range relationships…
OMG-Bench: A New Challenging Benchmark for Skeleton-based Online Micro Hand Gesture Recognition
Haochen Chang, Pengfei Ren, Buyuan Zhang +6
Online micro gesture recognition from hand skeletons is critical for VR/AR interaction but faces challenges due to limited public datasets and task-specific algorithms. Micro gestu…
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
Linzhi Wu, Xingyu Zhang, Hao Yuan +5
Audio-visual speech recognition (AVSR) typically improves recognition accuracy in noisy environments by integrating noise-immune visual cues with audio signals. Nevertheless, high-…
DBMIF: a deep balanced multimodal iterative fusion framework for air- and bone-conduction speech enhancement
Yilei Wu, Changyan Zheng, Xingyu Zhang +5
The performance of conventional speech enhancement systems degrades sharply in extremely low signal-to-noise ratio (SNR) environments where air-conduction (AC) microphones are over…
AFD-SLU: Adaptive Feature Distillation for Spoken Language Understanding
Yan Xie, Yibo Cui, Liang Xie +1
Spoken Language Understanding (SLU) is a core component of conversational systems, enabling machines to interpret user utterances. Despite its importance, developing effective SLU…