3 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.SD2025
Conditional Latent Diffusion-Based Speech Enhancement Via Dual Context Learning
Shengkui Zhao, Zexu Pan, Kun Zhou +3
Recently, the application of diffusion probabilistic models has advanced speech enhancement through generative approaches. However, existing diffusion-based methods have focused on…
cs.SD2025
HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution
Shengkui Zhao, Kun Zhou, Zexu Pan +3
The application of generative adversarial networks (GANs) has recently advanced speech super-resolution (SR) based on intermediate representations like mel-spectrograms. However, e…
cs.CL2025★ 3 cited
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Qian Chen, Yafeng Chen, Yanni Chen +33
Recent advancements in large language models (LLMs) and multimodal speech-text models have laid the groundwork for seamless voice interactions, enabling real-time, natural, and hum…