activity
20192022
most citedPHASEN: A Phase-and-Harmonics-Aware Speech Enhancement Network

26 citations · 38 across the 5 of their papers we have counts for

collaborators

5 papers

eess.AS2022

TridentSE: Guiding Speech Enhancement with 32 Global Tokens

Dacheng Yin, Zhiyuan Zhao, Chuanxin Tang +2

In this paper, we present TridentSE, a novel architecture for speech enhancement, which is capable of efficiently capturing both global information and local details. TridentSE mai…

cs.LG20224 cited

Retriever: Learning Content-Style Representation as a Token-Level Bipartite Graph

Dacheng Yin, Xuanchi Ren, Chong Luo +3

This paper addresses the unsupervised learning of content-style decomposed representation. We first give a definition of style and then model the content-style representation as a…

cs.SD2021

Zero-Shot Text-to-Speech for Text-Based Insertion in Audio Narration

Chuanxin Tang, Chong Luo, Zhiyuan Zhao +3

Given a piece of speech and its transcript text, text-based speech editing aims to generate speech that can be seamlessly inserted into the given speech by editing the transcript.…

cs.SD20218 cited

General-Purpose Speech Representation Learning through a Self-Supervised Multi-Granularity Framework

Yucheng Zhao, Dacheng Yin, Chong Luo +4

This paper presents a self-supervised learning framework, named MGF, for general-purpose speech representation learning. In the design of MGF, speech hierarchy is taken into consid…

cs.SD201926 cited

PHASEN: A Phase-and-Harmonics-Aware Speech Enhancement Network

Dacheng Yin, Chong Luo, Zhiwei Xiong +1

Time-frequency (T-F) domain masking is a mainstream approach for single-channel speech enhancement. Recently, focuses have been put to phase prediction in addition to amplitude pre…