5 papers
A Variational-Flow Analysis of Diffusion-Based Speech Enhancement under Noise-Power Mismatch
Shuubham Ojha
Diffusion-based speech enhancement architectures that pair a deterministic predictor with a learned score network, exhibit a sharp non-smooth transition (``kink'') in the SI-SDR de…
CaReCoS: A Spectrogram based Visual Benchmark for Cardiac, Respiratory and Cough Sounds
Harshit Rajgarhia, Shuubham Ojha, Akhil Pothanapalli +4
Medical acoustic signals such as respiratory sounds, cardiac auscultations, and cough audio carry rich diagnostic information, yet no existing benchmark evaluates multimodal reason…
Bridging Self-Supervised Learning and Speech Enhancement: A Wav2Vec2-Conditioned Framework
Shuubham Ojha, Carol Espy-Wilson
Diffusion models show potential for speech enhancement but lack linguistic guidance. We condition a diffusion-based model on wav2vec 2.0 features from noisy input, injected at the…
MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio
Harshit Rajgarhia, Shuubham Ojha, Asif Shaik +4
Medical audio data is difficult to collect due to privacy regulations and high annotation costs arising from domain expertise. Thus, existing benchmarks tend to underrepresent comp…
Reverse Attention for Lightweight Speech Enhancement on Edge Devices
Shuubham Ojha, Felix Gervits, Carol Espy-Wilson
This paper introduces a lightweight deep learning model for real-time speech enhancement, designed to operate efficiently on resource-constrained devices. The proposed model levera…