activity
20182023
most citedLearning Depth-Guided Convolutions for Monocular 3D Object Detection

30 citations · 101 across the 13 of their papers we have counts for

collaborators
Showing 2023Show all

9 papers · 1 filter

cs.CV20231 cited

Towards Free Data Selection with General-Purpose Models

Yichen Xie, Mingyu Ding, Masayoshi Tomizuka +1

A desirable data selection algorithm can efficiently choose the most informative samples to maximize the utility of limited annotation budgets. However, current approaches, represe…

cs.CV2023

Pre-training on Synthetic Driving Data for Trajectory Prediction

Yiheng Li, Seth Z. Zhao, Chenfeng Xu +5

Accumulating substantial volumes of real-world driving data proves pivotal in the realm of trajectory forecasting for autonomous driving. Given the heavy reliance of current trajec…

cs.CV2023

Physion++: Evaluating Physical Scene Understanding that Requires Online Inference of Different Physical Properties

Hsiao-Yu Tung, Mingyu Ding, Zhenfang Chen +6

General physical scene understanding requires more than simply localizing and recognizing objects -- it requires knowledge that objects can have different latent properties (e.g.,…

cs.RO2023

EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought

Yao Mu, Qinglong Zhang, Mengkang Hu +7

Embodied AI is a crucial frontier in robotics, capable of planning and executing action sequences for robots to accomplish long-horizon tasks in physical environments. In this work…

cs.CV2023

Quadric Representations for LiDAR Odometry, Mapping and Localization

Chao Xia, Chenfeng Xu, Patrick Rim +5

Current LiDAR odometry, mapping and localization methods leverage point-wise representations of 3D scenes and achieve high accuracy in autonomous driving tasks. However, the space-…

cs.LG2023

EC^2: Emergent Communication for Embodied Control

Yao Mu, Shunyu Yao, Mingyu Ding +2

Embodied control requires agents to leverage multi-modal pre-training to quickly learn how to act in new environments, where video demonstrations contain visual and motion details…