15 papers
wav2tok 2.0: Scalable Audio Tokenization Maintaining Explicit Pairwise Token Alignment for Efficient Audio Retrieval
Adhiraj Banerjee, Vipul Arora
Learning discrete speech representations that preserve similarity across variable-length utterances is central to query-by-example spoken term detection (QbE-STD). While wav2tok in…
PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization
Adhiraj Banerjee, Vipul Arora
Modern learning systems represent perceptual signals with continuous vectors, but comparison, retrieval, memory, alignment, and reasoning are often naturally symbolic. In language,…
Automatic Detection and Analysis of Singing Mistakes for Music Pedagogy
Sumit Kumar, Suraj Jaiswal, Parampreet Singh +1
The advancement of machine learning in audio analysis has opened new possibilities for technology-enhanced music education. This paper introduces a framework for automatic singing…
Learning to Discover: A Generalized Framework for Raga Identification without Forgetting
Parampreet Singh, Somya Kumar, Chaitanya Shailendra Nitawe +1
Raga identification in Indian Art Music (IAM) remains challenging due to the presence of numerous rarely performed Ragas that are not represented in available training datasets. Tr…
TORRCH: Tomographic reconstruction of the reionization of cosmic hydrogen with Ly emitters and non-Ly-selected galaxies
Soumak Maitra, Girish Kulkarni, Vipul Arora +5
Tomographic reconstruction of reionization is a long-sought goal. It would move the field beyond global summary statistics, such as the volume-averaged ionised fraction, to direct,…
Weakly Supervised Tabla Stroke Transcription via TI-SDRM: A Rhythm-Aware Lattice Rescoring Framework
Rahul Bapusaheb Kodag, Vipul Arora
Tabla Stroke Transcription (TST) is central to the analysis of rhythmic structure in Hindustani classical music, yet remains challenging due to complex rhythmic organization and th…