7 papers
Systematic Evaluation of Time-Frequency Features for Binaural Sound Source Localization
Davoud Shariat Panah, Alessandro Ragano, Dan Barry +2
This study presents a systematic evaluation of time-frequency feature design for binaural sound source localization (SSL), focusing on how feature selection influences model perfor…
Binaspect -- A Python Library for Binaural Audio Analysis, Visualization & Feature Generation
Dan Barry, Davoud Shariat Panah, Alessandro Ragano +2
We present Binaspect, an open-source Python library for binaural audio analysis, visualization, and feature generation. Binaspect generates interpretable "azimuth maps" by calculat…
BINAQUAL: A Full-Reference Objective Localization Similarity Metric for Binaural Audio
Davoud Shariat Panah, Dan Barry, Alessandro Ragano +2
Spatial audio enhances immersion in applications such as virtual reality, augmented reality, gaming, and cinema by creating a three-dimensional auditory experience. Ensuring the sp…
DM-Codec: Distilling Multimodal Representations for Speech Tokenization
Md Mubtasim Ahasan, Md Fahim, Tasnim Mohiuddin +6
Recent advancements in speech-language models have yielded significant improvements in speech tokenization and synthesis. However, effectively mapping the complex, multidimensional…
Binamix -- A Python Library for Generating Binaural Audio Datasets
Dan Barry, Davoud Shariat Panah, Alessandro Ragano +2
The increasing demand for spatial audio in applications such as virtual reality, immersive media, and spatial audio research necessitates robust solutions to generate binaural audi…
SCOREQ: Speech Quality Assessment with Contrastive Regression
Alessandro Ragano, Jan Skoglund, Andrew Hines
In this paper, we present SCOREQ, a novel approach for speech quality prediction. SCOREQ is a triplet loss function for contrastive regression that addresses the domain generalisat…