3 papers
cs.CV2022
Mix and Localize: Localizing Sound Sources in Mixtures
Xixi Hu, Ziyang Chen, Andrew Owens
We present a method for simultaneously localizing multiple sound sources within a visual scene. This task requires a model to both group a sound mixture into individual sources, an…
cs.RO2022
A Robotic Visual Grasping Design: Rethinking Convolution Neural Network with High-Resolutions
Zhangli Zhou, Shaochen Wang, Ziyang Chen +2
High-resolution representations are important for vision-based robotic grasping problems. Existing works generally encode the input images into low-resolution representations via s…
cs.CV2019
PSDNet and DPDNet: Efficient channel expansion, Depthwise-Pointwise-Depthwise Inverted Bottleneck Block
Guoqing Li, Meng Zhang, Qianru Zhang +7
In many real-time applications, the deployment of deep neural networks is constrained by high computational cost and efficient lightweight neural networks are widely concerned. In…