1 paper
Xixi Hu, Ziyang Chen, Andrew Owens
We present a method for simultaneously localizing multiple sound sources within a visual scene. This task requires a model to both group a sound mixture into individual sources, an…