activity
20212024
most citedPMVC: Data Augmentation-Based Prosody Modeling for Expressive Voice Conversion

15 citations · 50 across the 31 of their papers we have counts for

collaborators

33 papers

eess.IV202416 cited

Deep Learning Segmentation of Ascites on Abdominal CT Scans for Automatic Volume Quantification

Benjamin Hou, Sung-Won Lee, Jung-Min Lee +4

Purpose: To evaluate the performance of an automated deep learning method in detecting ascites and subsequently quantifying its volume in patients with liver cirrhosis and ovarian…

cs.CL2024

PFID: Privacy First Inference Delegation Framework for LLMs

Haoyan Yang, Zhitao Li, Yong Zhang +4

This paper introduces a novel privacy-preservation framework named PFID for LLMs that addresses critical privacy concerns by localizing user data through model sharding and singula…

cs.SD2024

EAD-VC: Enhancing Speech Auto-Disentanglement for Voice Conversion with IFUB Estimator and Joint Text-Guided Consistent Learning

Ziqi Liang, Jianzong Wang, Xulong Zhang +3

Using unsupervised learning to disentangle speech into content, rhythm, pitch, and timbre for voice conversion has become a hot research topic. Existing works generally take into a…

cs.SD2024

CONTUNER: Singing Voice Beautifying with Pitch and Expressiveness Condition

Jianzong Wang, Pengcheng Li, Xulong Zhang +2

Singing voice beautifying is a novel task that has application value in people's daily life, aiming to correct the pitch of the singing voice and improve the expressiveness without…

cs.AI2024

Medical Speech Symptoms Classification via Disentangled Representation

Jianzong Wang, Pengcheng Li, Xulong Zhang +2

Intent is defined for understanding spoken language in existing works. Both textual features and acoustic features involved in medical speech contain intent, which is important for…

cs.SD202413 cited

Retrieval-Augmented Audio Deepfake Detection

Zuheng Kang, Yayun He, Botao Zhao +4

With recent advances in speech synthesis including text-to-speech (TTS) and voice conversion (VC) systems enabling the generation of ultra-realistic audio deepfakes, there is growi…