2 citations · 2 across the 13 of their papers we have counts for
6 papers · 1 filter
BEST-STD2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection
Anup Singh, Vipul Arora, Kris Demuynck
Fast and accurate spoken content retrieval is vital for applications such as voice search. Query-by-Example Spoken Term Detection (STD) involves retrieving matching segments from a…
AudioNet: Supervised Deep Hashing for Retrieval of Similar Audio Events
Sagar Dutta, Vipul Arora
This work presents a supervised deep hashing method for retrieving similar audio events. The proposed method, named AudioNet, is a deep-learning-based system for efficient hashing…
H-QuEST: Accelerating Query-by-Example Spoken Term Detection with Hierarchical Indexing
Akanksha Singh, Yi-Ping Phoebe Chen, Vipul Arora
Query-by-example spoken term detection (QbE-STD) searches for matching words or phrases in an audio dataset using a sample spoken query. When annotated data is limited or unavailab…
Uncertainty Quantification in Melody Estimation using Histogram Representation
Kavya Ranjan Saxena, Vipul Arora
Confidence estimation can improve the reliability of melody estimation by indicating which predictions are likely incorrect. The existing classification-based approach provides con…
BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection
Anup Singh, Kris Demuynck, Vipul Arora
Spoken term detection (STD) is often hindered by reliance on frame-level features and the computationally intensive DTW-based template matching, limiting its practicality. To addre…
Interactive singing melody extraction based on active adaptation
Kavya Ranjan Saxena, Vipul Arora
Extraction of predominant pitch from polyphonic audio is one of the fundamental tasks in the field of music information retrieval and computational musicology. To accomplish this t…