activity
20202023
most citedAn Anchor-Free Detector for Continuous Speech Keyword Spotting

1 citations · 2 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2023

ARTV: Auto-Regressive Text-to-Video Generation with Diffusion Models

Wenming Weng, Ruoyu Feng, Yanhui Wang +10

We present ARTV, an efficient framework for auto-regressive video generation with diffusion models. Unlike existing methods that generate entire videos in one-s…

eess.AS2023

Filler Word Detection with Hard Category Mining and Inter-Category Focal Loss

Zhiyuan Zhao, Lijun Wu, Chuanxin Tang +3

Filler words like ``um" or ``uh" are common in spontaneous speech. It is desirable to automatically detect and remove them in recordings, as they affect the fluency, confidence, an…

eess.AS2022

TridentSE: Guiding Speech Enhancement with 32 Global Tokens

Dacheng Yin, Zhiyuan Zhao, Chuanxin Tang +2

In this paper, we present TridentSE, a novel architecture for speech enhancement, which is capable of efficiently capturing both global information and local details. TridentSE mai…

eess.AS2022★ 1 cited

An Anchor-Free Detector for Continuous Speech Keyword Spotting

Zhiyuan Zhao, Chuanxin Tang, Chengdong Yao +1

Continuous Speech Keyword Spotting (CSKWS) is a task to detect predefined keywords in a continuous speech. In this paper, we regard CSKWS as a one-dimensional object detection task…

eess.AS2022

RetrieverTTS: Modeling Decomposed Factors for Text-Based Speech Insertion

Dacheng Yin, Chuanxin Tang, Yanqing Liu +6

This paper proposes a new "decompose-and-edit" paradigm for the text-based speech insertion task that facilitates arbitrary-length speech insertion and even full sentence generatio…

cs.SD2021

Zero-Shot Text-to-Speech for Text-Based Insertion in Audio Narration

Chuanxin Tang, Chong Luo, Zhiyuan Zhao +3

Given a piece of speech and its transcript text, text-based speech editing aims to generate speech that can be seamlessly inserted into the given speech by editing the transcript.…