31 citations · 39 across the 4 of their papers we have counts for
4 papers
How to Listen? Rethinking Visual Sound Localization
Ho-Hsiang Wu, Magdalena Fuentes, Prem Seetharaman +1
Localizing visual sounds consists on locating the position of objects that emit sound within an image. It is a growing research area with potential applications in monitoring natur…
Exploring modality-agnostic representations for music classification
Ho-Hsiang Wu, Magdalena Fuentes, Juan P. Bello
Music information is often conveyed or recorded across multiple data modalities including but not limited to audio, images, text and scores. However, music information retrieval re…
Multi-Task Self-Supervised Pre-Training for Music Classification
Ho-Hsiang Wu, Chieh-Chi Kao, Qingming Tang +4
Deep learning is very data hungry, and supervised learning especially requires massive labeled data to work well. Machine listening research often suffers from limited labeled data…
SONYC-UST-V2: An Urban Sound Tagging Dataset with Spatiotemporal Context
Mark Cartwright, Jason Cramer, Ana Elisa Mendez Mendez +9
We present SONYC-UST-V2, a dataset for urban sound tagging with spatiotemporal information. This dataset is aimed for the development and evaluation of machine listening systems fo…