137 citations · 206 across the 12 of their papers we have counts for
11 papers · 1 filter
E^2VPT: An Effective and Efficient Approach for Visual Prompt Tuning
Cheng Han, Qifan Wang, Yiming Cui +4
As the size of transformer-based models continues to grow, fine-tuning these large-scale pretrained vision models for new tasks has become increasingly parameter-intensive. Paramet…
Cascaded Human-Object Interaction Recognition
Tianfei Zhou, Wenguan Wang, Siyuan Qi +2
Rapid progress has been witnessed for human-object interaction (HOI) recognition, but most existing models are confined to single-stage reasoning pipelines. Considering the intrins…
Learning Compositional Neural Information Fusion for Human Parsing
Wenguan Wang, Zhijie Zhang, Siyuan Qi +3
This work proposes to combine neural networks with the compositional hierarchy of human bodies for efficient and complete human parsing. We formulate the approach as a neural infor…
PerspectiveNet: 3D Object Detection from a Single RGB Image via Perspective Points
Siyuan Huang, Yixin Chen, Tao Yuan +3
Detecting 3D objects from a single RGB image is intrinsically ambiguous, thus requiring appropriate prior knowledge and intermediate representations as constraints to reduce the un…
Holistic++ Scene Understanding: Single-view 3D Holistic Scene Parsing and Human Pose Estimation with Human-Object Interaction and Physical Commonsense
Yixin Chen, Siyuan Huang, Tao Yuan +3
We propose a new 3D holistic++ scene understanding problem, which jointly tackles two tasks from a single-view image: (i) holistic scene parsing and reconstruction---3D estimations…
Reasoning Visual Dialogs with Structural and Partial Observations
Zilong Zheng, Wenguan Wang, Siyuan Qi +1
We propose a novel model to address the task of Visual Dialog which exhibits complex dialog structures. To obtain a reasonable answer based on the current question and the dialog h…