activity
20182022
most citedLearning Depth-Guided Convolutions for Monocular 3D Object Detection

30 citations · 94 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

13 papers · 1 filter

cs.CV202214 cited

ComPhy: Compositional Physical Reasoning of Objects and Events from Videos

Zhenfang Chen, Kexin Yi, Yunzhu Li +4

Objects' motions in nature are governed by complex interactions and their properties. While some properties, such as shape and material, can be identified via the object's visual a…

cs.CV202211 cited

DaViT: Dual Attention Vision Transformers

Mingyu Ding, Bin Xiao, Noel Codella +3

In this work, we introduce Dual Attention Vision Transformers (DaViT), a simple yet effective vision transformer architecture that is able to capture global context while maintaini…

cs.CV202110 cited

Dynamic Visual Reasoning by Learning Differentiable Physics Models from Video and Language

Mingyu Ding, Zhenfang Chen, Tao Du +3

In this work, we propose a unified framework, called Visual Reasoning with Differ-entiable Physics (VRDP), that can jointly learn visual concepts and infer physics models of object…

cs.CV2021

PolarMask++: Enhanced Polar Representation for Single-Shot Instance Segmentation and Beyond

Enze Xie, Wenhai Wang, Mingyu Ding +2

Reducing the complexity of the pipeline of instance segmentation is crucial for real-world applications. This work addresses this issue by introducing an anchor-box free and single…

cs.CV202010 cited

Dense Hybrid Recurrent Multi-view Stereo Net with Dynamic Consistency Checking

Jianfeng Yan, Zizhuang Wei, Hongwei Yi +5

In this paper, we propose an efficient and effective dense hybrid recurrent multi-view stereo net with dynamic consistency checking, namely HC-RMVSNet, for accurate dense po…

cs.CV2020

Domain-Adaptive Few-Shot Learning

An Zhao, Mingyu Ding, Zhiwu Lu +5

Existing few-shot learning (FSL) methods make the implicit assumption that the few target class samples are from the same domain as the source class samples. However, in practice t…