Publications (25)
MessyTable: Instance Association in Multiple Camera Views
Zhongang Cai, Junzhe Zhang, Daxuan Ren +5
We present an interesting and challenging dataset that features a large number of scenes with messy tables captured from multiple camera views. Each scene in this dataset is highly…
Differentiable Particles for General-Purpose Deformable Object Manipulation
Siwei Chen, Yiqing Xu, Cunjun Yu +2
Deformable object manipulation is a long-standing challenge in robotics. While existing approaches often focus narrowly on a specific type of object, we seek a general-purpose algo…
What Truly Matters in Trajectory Prediction for Autonomous Driving?
Phong Tran, Haoran Wu, Cunjun Yu +3
Trajectory prediction plays a vital role in the performance of autonomous driving systems, and prediction accuracy, such as average displacement error (ADE) or final displacement e…
Learning an Efficient Network for Large-Scale Hierarchical Object Detection with Data Imbalance: 3rd Place Solution to Open Images Challenge 2019
Xingyuan Bu, Junran Peng, Changbao Wang +2
This report details our solution to the Google AI Open Images Challenge 2019 Object Detection Track. Based on our detailed analysis on the Open Images dataset, it is found that the…
INVIGORATE: Interactive Visual Grounding and Grasping in Clutter
Hanbo Zhang, Yunfan Lu, Cunjun Yu +3
This paper presents INVIGORATE, a robot system that interacts with human through natural language and grasps a specified object in clutter. The objects may occlude, obstruct, or ev…
Vision-Language Foundation Models as Effective Robot Imitators
Xinghang Li, Minghuan Liu, Hanbo Zhang +9
Recent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipul…
Leveraging Temporal Information for 3D Detection and Domain Adaptation
Cunjun Yu, Zhongang Cai, Daxuan Ren +1
Ever since the prevalent use of the LiDARs in autonomous driving, tremendous improvements have been made to the learning on the point clouds. However, recent progress largely focus…
Robi Butler: Multimodal Remote Interaction with a Household Robot Assistant
Anxing Xiao, Nuwan Janaka, Tianrun Hu +4
Imagine a future when we can Zoom-call a robot to manage household chores remotely. This work takes one step in this direction. Robi Butler is a new household robot assistant that…
"Stack It Up!": 3D Stable Structure Generation from 2D Hand-drawn Sketch
Yiqing Xu, Linfeng Li, Cunjun Yu +1
Imagine a child sketching the Eiffel Tower and asking a robot to bring it to life. Today's robot manipulation systems can't act on such sketches directly-they require precise 3D bl…
Balanced MSE for Imbalanced Visual Regression
Jiawei Ren, Mingyuan Zhang, Cunjun Yu +1
Data imbalance exists ubiquitously in real-world visual regressions, e.g., age estimation and pose estimation, hurting the model's generalizability and fairness. Thus, imbalanced r…
DaXBench: Benchmarking Deformable Object Manipulation with Differentiable Physics
Siwei Chen, Yiqing Xu, Cunjun Yu +4
Deformable Object Manipulation (DOM) is of significant importance to both daily and industrial applications. Recent successes in differentiable physics simulators allow learning al…
DistillNeRF: Perceiving 3D Scenes from Single-Glance Images by Distilling Neural Fields and Foundation Model Features
Letian Wang, Seung Wook Kim, Jiawei Yang +7
We propose DistillNeRF, a self-supervised learning framework addressing the challenge of understanding 3D environments from limited 2D observations in outdoor autonomous driving sc…
CLASP: General-Purpose Clothes Manipulation with Semantic Keypoints
Yuhong Deng, Chao Tang, Cunjun Yu +2
Clothes manipulation, such as folding or hanging, is a critical capability for home service robots. Despite recent advances, most existing methods remain limited to specific clothe…
InsActor: Instruction-driven Physics-based Characters
Jiawei Ren, Mingyuan Zhang, Cunjun Yu +3
Generating animation of physics-based characters with intuitive control has long been a desirable task with numerous applications. However, generating physically simulated animatio…
CANINE: Coaching Visually Impaired Users for Interactive Navigation with a Robot Guide Dog
Cunjun Yu, Zishuo Wang, Anxing Xiao +2
Robot guide dogs offer navigation assistance that greatly expands the independent mobility of the visually impaired, but their effective use requires subtle human-robot coordinatio…
COACH: Cooperative Robot Teaching
Cunjun Yu, Yiqing Xu, Linfeng Li +1
Knowledge and skills can transfer from human teachers to human students. However, such direct transfer is often not scalable for physical tasks, as they require one-to-one interact…
DiffMimic: Efficient Motion Mimicking with Differentiable Physics
Jiawei Ren, Cunjun Yu, Siwei Chen +3
Motion mimicking is a foundational task in physics-based character animation. However, most existing motion mimicking methods are built upon reinforcement learning (RL) and suffer…
Balanced Meta-Softmax for Long-Tailed Visual Recognition
Jiawei Ren, Cunjun Yu, Shunan Sheng +4
Deep classifiers have achieved great success in visual recognition. However, real-world data is long-tailed by nature, leading to the mismatch between training and testing distribu…
3D Convolution on RGB-D Point Clouds for Accurate Model-free Object Pose Estimation
Zhongang Cai, Cunjun Yu, Quang-Cuong Pham
The conventional pose estimation of a 3D object usually requires the knowledge of the 3D model of the object. Even with the recent development in convolutional neural networks (CNN…
GSON: A Group-based Social Navigation Framework with Large Multimodal Model
Shangyi Luo, Peng Sun, Ji Zhu +4
With the increasing presence of service robots and autonomous vehicles in human environments, navigation systems need to evolve beyond simple destination reach to incorporate socia…
TeachingBot: Robot Teacher for Human Handwriting
Zhimin Hou, Cunjun Yu, David Hsu +1
Teaching and learning physical skills often require one-on-one interaction, making it difficult to scale up, as there are not enough human teachers. Robots offer an attractive alte…
Leveraging Localization for Multi-camera Association
Zhongang Cai, Cunjun Yu, Junzhe Zhang +2
We present McAssoc, a deep learning approach to the as-sociation of detection bounding boxes in different views ofa multi-camera system. The vast majority of the academiahas been d…
Balanced Activation for Long-tailed Visual Recognition
Jiawei Ren, Cunjun Yu, Zhongang Cai +1
Deep classifiers have achieved great success in visual recognition. However, real-world data is long-tailed by nature, leading to the mismatch between training and testing distribu…
Spatio-Temporal Graph Transformer Networks for Pedestrian Trajectory Prediction
Cunjun Yu, Xiao Ma, Jiawei Ren +2
Understanding crowd motion dynamics is critical to real-world applications, e.g., surveillance systems and autonomous driving. This is challenging because it requires effectively m…
Siamese Convolutional Neural Network for Sub-millimeter-accurate Camera Pose Estimation and Visual Servoing
Cunjun Yu, Zhongang Cai, Hung Pham +1
Visual Servoing (VS), where images taken from a camera typically attached to the robot end-effector are used to guide the robot motions, is an important technique to tackle robotic…