9 papers
CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents
Adhiraj Banerjee, Vipul Arora
Text-guided sound separation enables flexible audio editing, assistive listening, and open-domain source extraction, but systems such as AudioSep remain too expensive for low-laten…
BEST-STD2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection
Anup Singh, Vipul Arora, Kris Demuynck
Fast and accurate spoken content retrieval is vital for applications such as voice search. Query-by-Example Spoken Term Detection (STD) involves retrieving matching segments from a…
Uncertainty Quantification in Melody Estimation using Histogram Representation
Kavya Ranjan Saxena, Vipul Arora
Confidence estimation can improve the reliability of melody estimation by indicating which predictions are likely incorrect. The existing classification-based approach provides con…
AudioNet: Supervised Deep Hashing for Retrieval of Similar Audio Events
Sagar Dutta, Vipul Arora
This work presents a supervised deep hashing method for retrieving similar audio events. The proposed method, named AudioNet, is a deep-learning-based system for efficient hashing…
SyncNet: correlating objective for time delay estimation in audio signals
Akshay Raina, Vipul Arora
This study addresses the task of performing robust and reliable time-delay estimation in signals in noisy and reverberating environments. In contrast to the popular signal processi…
H-QuEST: Accelerating Query-by-Example Spoken Term Detection with Hierarchical Indexing
Akanksha Singh, Yi-Ping Phoebe Chen, Vipul Arora
Query-by-example spoken term detection (QbE-STD) searches for matching words or phrases in an audio dataset using a sample spoken query. When annotated data is limited or unavailab…