2 citations · 2 across the 1 of their papers we have counts for
2 papers
cs.CL2026
Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning
Hayato Futami, Hassan Shahmohammadi, Tushar Dhyani +5
Speech-to-speech translation (S2ST) has advanced significantly with speech LLMs, offering the potential for joint optimization and preserving non-linguistic information. However, t…
cs.SD2024★ 2 cited
High-Resolution Speech Restoration with Latent Diffusion Model
Tushar Dhyani, Florian Lux, Michele Mancusi +3
Traditional speech enhancement methods often oversimplify the task of restoration by focusing on a single type of distortion. Generative models that handle multiple distortions fre…