activity
20242026
most citedLLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems

2 citations · 3 across the 13 of their papers we have counts for

collaborators
Showing 2024Show all

7 papers · 1 filter

eess.AS2024

Video-to-Audio Generation with Fine-grained Temporal Semantics

Yuchen Hu, Yu Gu, Chenxing Li +2

With recent advances of AIGC, video generation have gained a surge of research interest in both academia and industry (e.g., Sora). However, it remains a challenge to produce tempo…

cs.SD2024

STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment

Yong Ren, Chenxing Li, Manjie Xu +4

Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmon…

eess.AS2024

SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis

Helin Wang, Meng Yu, Jiarui Hai +5

In this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot textbased speech editing and text-to-speech synthesis. S…

eess.AS2024

LCM-SVC: Latent Diffusion Model Based Singing Voice Conversion with Inference Acceleration via Latent Consistency Distillation

Shihao Chen, Yu Gu, Jianwei Cui +3

Any-to-any singing voice conversion (SVC) aims to transfer a target singer's timbre to other songs using a short voice sample. However many diffusion model based any-to-any SVC met…

cs.SD20241 cited

Video-to-Audio Generation with Hidden Alignment

Manjie Xu, Chenxing Li, Xinyi Tu +5

Generating semantically and temporally aligned audio content in accordance with video input has become a focal point for researchers, particularly following the remarkable breakthr…

eess.AS2024

LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance

Shihao Chen, Yu Gu, Jie Zhang +4

Any-to-any singing voice conversion (SVC) is an interesting audio editing technique, aiming to convert the singing voice of one singer into that of another, given only a few second…