7 papers · 1 filter
Can Hierarchical Cross-Modal Fusion Predict Human Perception of AI Dubbed Content?
Ashwini Dasare, Nirmesh Shah, Ashishkumar Gudmalwar +1
Evaluating AI generated dubbed content is inherently multi-dimensional, shaped by synchronization, intelligibility, speaker consistency, emotional alignment, and semantic context.…
Gesture2Speech: How Far Can Hand Movements Shape Expressive Speech?
Lokesh Kumar, Nirmesh Shah, Ashishkumar P. Gudmalwar +1
Human communication seamlessly integrates speech and bodily motion, where hand gestures naturally complement vocal prosody to express intent, emotion, and emphasis. While recent te…
Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-resource Speech Recognition
Aditya Srinivas Menon, Kumud Tripathi, Raj Gohil +1
Self-supervised learning (SSL) has advanced speech processing but suffers from quadratic complexity due to self-attention. To address this, SummaryMixing (SM) has been proposed as…
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
Ishan D. Biyani, Nirmesh J. Shah, Ashishkumar P. Gudmalwar +2
Speech time reversal refers to the process of reversing the entire speech signal in time, causing it to play backward. Such signals are completely unintelligible since the fundamen…
EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion
Ashishkumar Gudmalwar, Ishan D. Biyani, Nirmesh Shah +2
The Emotional Voice Conversion (EVC) aims to convert the discrete emotional state from the source emotion to the target for a given speech utterance while preserving linguistic con…
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
Neha Sahipjohn, Ashishkumar Gudmalwar, Nirmesh Shah +2
Audio-visual alignment after dubbing is a challenging research problem. To this end, we propose a novel method, DubWise Multi-modal Large Language Model (LLM)-based Text-to-Speech…