9 citations · 17 across the 5 of their papers we have counts for
5 papers
iQuery: Instruments as Queries for Audio-Visual Sound Separation
Jiaben Chen, Renrui Zhang, Dongze Lian +3
Current audio-visual separation methods share a standard architecture design where an audio encoder-decoder network is fused with visual encoding features at the encoder bottleneck…
TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting
Huazhang Hu, Sixun Dong, Yiqun Zhao +3
Counting repetitive actions are widely seen in human activities such as physical exercise. Existing methods focus on performing repetitive action counting in short videos, which is…
Look Before You Leap: Learning Landmark Features for One-Stage Visual Grounding
Binbin Huang, Dongze Lian, Weixin Luo +1
An LBYL (`Look Before You Leap') Network is proposed for end-to-end trainable one-stage visual grounding. The idea behind LBYL-Net is intuitive and straightforward: we follow a lan…
Believe It or Not, We Know What You Are Looking at!
Dongze Lian, Zehao Yu, Shenghua Gao
By borrowing the wisdom of human in gaze following, we propose a two-stage solution for gaze point prediction of the target persons in a scene. Specifically, in the first stage, bo…
Single-Image Piece-wise Planar 3D Reconstruction via Associative Embedding
Zehao Yu, Jia Zheng, Dongze Lian +2
Single-image piece-wise planar 3D reconstruction aims to simultaneously segment plane instances and recover 3D plane parameters from an image. Most recent approaches leverage convo…