activity
20242026
collaborators

9 papers

cs.SD2026

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents

Adhiraj Banerjee, Vipul Arora

Text-guided sound separation enables flexible audio editing, assistive listening, and open-domain source extraction, but systems such as AudioSep remain too expensive for low-laten…

eess.AS2026

BEST-STD2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection

Anup Singh, Vipul Arora, Kris Demuynck

Fast and accurate spoken content retrieval is vital for applications such as voice search. Query-by-Example Spoken Term Detection (STD) involves retrieving matching segments from a…

eess.AS2025

Uncertainty Quantification in Melody Estimation using Histogram Representation

Kavya Ranjan Saxena, Vipul Arora

Confidence estimation can improve the reliability of melody estimation by indicating which predictions are likely incorrect. The existing classification-based approach provides con…

eess.AS2025

AudioNet: Supervised Deep Hashing for Retrieval of Similar Audio Events

Sagar Dutta, Vipul Arora

This work presents a supervised deep hashing method for retrieving similar audio events. The proposed method, named AudioNet, is a deep-learning-based system for efficient hashing…

eess.AS2025

SyncNet: correlating objective for time delay estimation in audio signals

Akshay Raina, Vipul Arora

This study addresses the task of performing robust and reliable time-delay estimation in signals in noisy and reverberating environments. In contrast to the popular signal processi…

eess.AS2025

H-QuEST: Accelerating Query-by-Example Spoken Term Detection with Hierarchical Indexing

Akanksha Singh, Yi-Ping Phoebe Chen, Vipul Arora

Query-by-example spoken term detection (QbE-STD) searches for matching words or phrases in an audio dataset using a sample spoken query. When annotated data is limited or unavailab…