1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2025
Video Forgery Detection with Optical Flow Residuals and Spatial-Temporal Consistency
Xi Xue, Kunio Suzuki, Nabarun Goswami +1
The rapid advancement of diffusion-based video generation models has led to increasingly realistic synthetic content, presenting new challenges for video forgery detection. Existin…
cs.SD2025
FUSE: Universal Speech Enhancement using Multi-Stage Fusion of Sparse Compression and Token Generation Models for the URGENT 2025 Challenge
Nabarun Goswami, Tatsuya Harada
We propose a multi-stage framework for universal speech enhancement, designed for the Interspeech 2025 URGENT Challenge. Our system first employs a Sparse Compression Network to ro…
eess.AS2022★ 1 cited
SATTS: Speaker Attractor Text to Speech, Learning to Speak by Learning to Separate
Nabarun Goswami, Tatsuya Harada
The mapping of text to speech (TTS) is non-deterministic, letters may be pronounced differently based on context, or phonemes can vary depending on various physiological and stylis…