31 citations · 35 across the 4 of their papers we have counts for
4 papers
How to Listen? Rethinking Visual Sound Localization
Ho-Hsiang Wu, Magdalena Fuentes, Prem Seetharaman +1
Localizing visual sounds consists on locating the position of objects that emit sound within an image. It is a growing research area with potential applications in monitoring natur…
Soundata: A Python library for reproducible use of audio datasets
Magdalena Fuentes, Justin Salamon, Pablo Zinemanas +6
Soundata is a Python library for loading and working with audio datasets in a standardized way, removing the need for writing custom loaders in every project, and improving reprodu…
Exploring modality-agnostic representations for music classification
Ho-Hsiang Wu, Magdalena Fuentes, Juan P. Bello
Music information is often conveyed or recorded across multiple data modalities including but not limited to audio, images, text and scores. However, music information retrieval re…
SONYC-UST-V2: An Urban Sound Tagging Dataset with Spatiotemporal Context
Mark Cartwright, Jason Cramer, Ana Elisa Mendez Mendez +9
We present SONYC-UST-V2, a dataset for urban sound tagging with spatiotemporal information. This dataset is aimed for the development and evaluation of machine listening systems fo…