88 citations · 204 across the 28 of their papers we have counts for
6 papers · 1 filter
Improving weakly supervised sound event detection with self-supervised auxiliary tasks
Soham Deshmukh, Bhiksha Raj, Rita Singh
While multitask and transfer learning has shown to improve the performance of neural networks in limited data settings, they require pretraining of the model on large datasets befo…
Multi-Task Learning for Interpretable Weakly Labelled Sound Event Detection
Soham Deshmukh, Bhiksha Raj, Rita Singh
Weakly Labelled learning has garnered lot of attention in recent years due to its potential to scale Sound Event Detection (SED) and is formulated as Multiple Instance Learning (MI…
Exploring Optimal DNN Architecture for End-to-End Beamformers Based on Time-frequency References
Yuichiro Koyama, Bhiksha Raj
Acoustic beamformers have been widely used to enhance audio signals. Currently, the best methods are the deep neural network (DNN)-powered variants of the generalized eigenvalue an…
Efficient Integration of Multi-channel Information for Speaker-independent Speech Separation
Yuichiro Koyama, Oluwafemi Azeez, Bhiksha Raj
Although deep-learning-based methods have markedly improved the performance of speech separation over the past few years, it remains an open question how to integrate multi-channel…
Exploring the Best Loss Function for DNN-Based Low-latency Speech Enhancement with Temporal Convolutional Networks
Yuichiro Koyama, Tyler Vuong, Stefan Uhlich +1
Recently, deep neural networks (DNNs) have been successfully used for speech enhancement, and DNN-based speech enhancement is becoming an attractive research area. While time-frequ…
The phonetic bases of vocal expressed emotion: natural versus acted
Hira Dhamyal, Shahan Ali Memon, Bhiksha Raj +1
Can vocal emotions be emulated? This question has been a recurrent concern of the speech community, and has also been vigorously investigated. It has been fueled further by its lin…