3 papers
cs.SD2025
Advancing NAM-to-Speech Conversion with Novel Methods and the MultiNAM Dataset
Neil Shah, Shirish Karande, Vineet Gandhi
Current Non-Audible Murmur (NAM)-to-speech techniques rely on voice cloning to simulate ground-truth speech from paired whispers. However, the simulated speech often lacks intellig…
cs.SD2025
MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI
Neil Shah, Ayan Kashyap, Shirish Karande +1
Previous real-time MRI (rtMRI)-based speech synthesis models depend heavily on noisy ground-truth speech. Applying loss directly over ground truth mel-spectrograms entangles speech…
cs.SD2024
Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models
Neil Shah, Shirish Karande, Vineet Gandhi
We propose a novel approach to significantly improve the intelligibility in the Non-Audible Murmur (NAM)-to-speech conversion task, leveraging self-supervision and sequence-to-sequ…