4 papers
Learning neural audio features without supervision
Sarthak Yadav, Neil Zeghidour
Deep audio classification, traditionally cast as training a deep neural network on top of mel-filterbanks in a supervised fashion, has recently benefited from two independent lines…
GISE-51: A scalable isolated sound events dataset
Sarthak Yadav, Mary Ellen Foster
Most of the existing isolated sound event datasets comprise a small number of sound event classes, usually 10 to 15, restricted to a small domain, such as domestic and urban sound…
End-to-End Bengali Speech Recognition
Sayan Mandal, Sarthak Yadav, Atul Rai
Bengali is a prominent language of the Indian subcontinent. However, while many state-of-the-art acoustic models exist for prominent languages spoken in the region, research and re…
Frequency and temporal convolutional attention for text-independent speaker recognition
Sarthak Yadav, Atul Rai
Majority of the recent approaches for text-independent speaker recognition apply attention or similar techniques for aggregation of frame-level feature descriptors generated by a d…