7 papers
ECHOv2: Two-Level Band-Splitting Representation Learning for Anomalous Sound Detection
Yucong Zhang, Juan Liu, Ming Li
Machine anomalous sound detection (ASD) requires robust audio representations capable of capturing subtle deviations in machine sounds under limited supervision. Existing pre-train…
Multi-View Based Audio Visual Target Speaker Extraction
Peijun Yang, Zhan Jin, Juan Liu +1
Audio-Visual Target Speaker Extraction (AVTSE) aims to separate a target speaker's voice from a mixed audio signal using the corresponding visual cues. While most existing AVTSE me…
Language-Invariant Multilingual Speaker Verification for the TidyVoice 2026 Challenge
Ze Li, Xiaoxiao Miao, Juan Liu +1
Multilingual speaker verification (SV) remains challenging due to limited cross-lingual data and language-dependent information in speaker embeddings. This paper presents a languag…
Toward Multimodal Industrial Fault Analysis: A Single-Speed Chain Conveyor Dataset with Audio and Vibration Signals
Zhang Chen, Yucong Zhang, Xiaoxiao Miao +1
We introduce a multimodal industrial fault analysis dataset collected from a single-speed chain conveyor (SSCC) system, targeting system-level fault detection in production lines.…
ECHO: Frequency-aware Hierarchical Encoding for Variable-length Signals
Yucong Zhang, Juan Liu, Ming Li
Pre-trained foundation models have demonstrated remarkable success in audio, vision and language, yet their potential for general machine signal modeling with arbitrary sampling ra…
Multimodal Laryngoscopic Video Analysis for Assisted Diagnosis of Vocal Fold Paralysis
Yucong Zhang, Xin Zou, Jinshan Yang +4
This paper presents the Multimodal Laryngoscopic Video Analyzing System (MLVAS), a novel system that leverages both audio and video data to automatically extract key video segments…