3 papers
eess.AS2023
Filler Word Detection with Hard Category Mining and Inter-Category Focal Loss
Zhiyuan Zhao, Lijun Wu, Chuanxin Tang +3
Filler words like ``um" or ``uh" are common in spontaneous speech. It is desirable to automatically detect and remove them in recordings, as they affect the fluency, confidence, an…
cs.CV2023
Streaming Video Model
Yucheng Zhao, Chong Luo, Chuanxin Tang +3
Video understanding tasks have traditionally been modeled by two separate architectures, specially tailored for two distinct tasks. Sequence-based video tasks, such as action recog…
eess.AS2022
RetrieverTTS: Modeling Decomposed Factors for Text-Based Speech Insertion
Dacheng Yin, Chuanxin Tang, Yanqing Liu +6
This paper proposes a new "decompose-and-edit" paradigm for the text-based speech insertion task that facilitates arbitrary-length speech insertion and even full sentence generatio…