activity
20172023
most citedTowards Realistic Visual Dubbing with Heterogeneous Sources

34 citations · 114 across the 28 of their papers we have counts for

collaborators
Showing cs.SDShow all

9 papers · 1 filter

cs.SD2022

Token-level Speaker Change Detection Using Speaker Difference and Speech Content via Continuous Integrate-and-fire

Zhiyun Fan, Zhenlin Liang, Linhao Dong +6

In multi-talker scenarios such as meetings and conversations, speech processing systems are usually required to segment the audio and then transcribe each segmentation. These two s…

cs.SD2022

The Volcspeech system for the ICASSP 2022 multi-channel multi-party meeting transcription challenge

Chen Shen, Yi Liu, Wenzhi Fan +6

This paper describes our submission to ICASSP 2022 Multi-channel Multi-party Meeting Transcription (M2MeT) Challenge. For Track 1, we propose several approaches to empower the clus…

cs.SD20221 cited

HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection

Ke Chen, Xingjian Du, Bilei Zhu +3

Audio classification is an important task of mapping audio samples into their corresponding labels. Recently, the transformer model with self-attention mechanisms has been adopted…

cs.SD20212 cited

Towards High-fidelity Singing Voice Conversion with Acoustic Reference and Contrastive Predictive Coding

Chao Wang, Zhonghao Li, Benlai Tang +4

Recently, phonetic posteriorgrams (PPGs) based methods have been quite popular in non-parallel singing voice conversion systems. However, due to the lack of acoustic information in…

cs.SD2021

Attention-based cross-modal fusion for audio-visual voice activity detection in musical video streams

Yuanbo Hou, Zhesong Yu, Xia Liang +4

Many previous audio-visual voice-related works focus on speech, ignoring the singing voice in the growing number of musical video streams on the Internet. For processing diverse mu…

cs.SD2020

Rule-embedded network for audio-visual voice activity detection in live musical video streams

Yuanbo Hou, Yi Deng, Bilei Zhu +2

Detecting anchor's voice in live musical streams is an important preprocessing for music and speech signal processing. Existing approaches to voice activity detection (VAD) primari…