papers

Publications (25)

cs.CV2020

MessyTable: Instance Association in Multiple Camera Views

Zhongang Cai, Junzhe Zhang, Daxuan Ren +5

We present an interesting and challenging dataset that features a large number of scenes with messy tables captured from multiple camera views. Each scene in this dataset is highly…

cs.RO2024

Differentiable Particles for General-Purpose Deformable Object Manipulation

Siwei Chen, Yiqing Xu, Cunjun Yu +2

Deformable object manipulation is a long-standing challenge in robotics. While existing approaches often focus narrowly on a specific type of object, we seek a general-purpose algo…

cs.RO2023

What Truly Matters in Trajectory Prediction for Autonomous Driving?

Phong Tran, Haoran Wu, Cunjun Yu +3

Trajectory prediction plays a vital role in the performance of autonomous driving systems, and prediction accuracy, such as average displacement error (ADE) or final displacement e…

cs.CV2019

Learning an Efficient Network for Large-Scale Hierarchical Object Detection with Data Imbalance: 3rd Place Solution to Open Images Challenge 2019

Xingyuan Bu, Junran Peng, Changbao Wang +2

This report details our solution to the Google AI Open Images Challenge 2019 Object Detection Track. Based on our detailed analysis on the Open Images dataset, it is found that the…

cs.RO2024

INVIGORATE: Interactive Visual Grounding and Grasping in Clutter

Hanbo Zhang, Yunfan Lu, Cunjun Yu +3

This paper presents INVIGORATE, a robot system that interacts with human through natural language and grasps a specified object in clutter. The objects may occlude, obstruct, or ev…

cs.RO2024

Vision-Language Foundation Models as Effective Robot Imitators

Xinghang Li, Minghuan Liu, Hanbo Zhang +9

Recent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipul…

cs.CV2020

Leveraging Temporal Information for 3D Detection and Domain Adaptation

Cunjun Yu, Zhongang Cai, Daxuan Ren +1

Ever since the prevalent use of the LiDARs in autonomous driving, tremendous improvements have been made to the learning on the point clouds. However, recent progress largely focus…

cs.RO2025

Robi Butler: Multimodal Remote Interaction with a Household Robot Assistant

Anxing Xiao, Nuwan Janaka, Tianrun Hu +4

Imagine a future when we can Zoom-call a robot to manage household chores remotely. This work takes one step in this direction. Robi Butler is a new household robot assistant that…

cs.AI2025

"Stack It Up!": 3D Stable Structure Generation from 2D Hand-drawn Sketch

Yiqing Xu, Linfeng Li, Cunjun Yu +1

Imagine a child sketching the Eiffel Tower and asking a robot to bring it to life. Today's robot manipulation systems can't act on such sketches directly-they require precise 3D bl…

cs.CV2022

Balanced MSE for Imbalanced Visual Regression

Jiawei Ren, Mingyuan Zhang, Cunjun Yu +1

Data imbalance exists ubiquitously in real-world visual regressions, e.g., age estimation and pose estimation, hurting the model's generalizability and fairness. Thus, imbalanced r…

cs.RO2023

DaXBench: Benchmarking Deformable Object Manipulation with Differentiable Physics

Siwei Chen, Yiqing Xu, Cunjun Yu +4

Deformable Object Manipulation (DOM) is of significant importance to both daily and industrial applications. Recent successes in differentiable physics simulators allow learning al…

cs.CV2024

DistillNeRF: Perceiving 3D Scenes from Single-Glance Images by Distilling Neural Fields and Foundation Model Features

Letian Wang, Seung Wook Kim, Jiawei Yang +7

We propose DistillNeRF, a self-supervised learning framework addressing the challenge of understanding 3D environments from limited 2D observations in outdoor autonomous driving sc…

cs.RO2025

CLASP: General-Purpose Clothes Manipulation with Semantic Keypoints

Yuhong Deng, Chao Tang, Cunjun Yu +2

Clothes manipulation, such as folding or hanging, is a critical capability for home service robots. Despite recent advances, most existing methods remain limited to specific clothe…

cs.CV2023

InsActor: Instruction-driven Physics-based Characters

Jiawei Ren, Mingyuan Zhang, Cunjun Yu +3

Generating animation of physics-based characters with intuitive control has long been a desirable task with numerous applications. However, generating physically simulated animatio…

cs.RO2026

CANINE: Coaching Visually Impaired Users for Interactive Navigation with a Robot Guide Dog

Cunjun Yu, Zishuo Wang, Anxing Xiao +2

Robot guide dogs offer navigation assistance that greatly expands the independent mobility of the visually impaired, but their effective use requires subtle human-robot coordinatio…

cs.RO2023

COACH: Cooperative Robot Teaching

Cunjun Yu, Yiqing Xu, Linfeng Li +1

Knowledge and skills can transfer from human teachers to human students. However, such direct transfer is often not scalable for physical tasks, as they require one-to-one interact…

cs.CV2023

DiffMimic: Efficient Motion Mimicking with Differentiable Physics

Jiawei Ren, Cunjun Yu, Siwei Chen +3

Motion mimicking is a foundational task in physics-based character animation. However, most existing motion mimicking methods are built upon reinforcement learning (RL) and suffer…

cs.LG2020

Balanced Meta-Softmax for Long-Tailed Visual Recognition

Jiawei Ren, Cunjun Yu, Shunan Sheng +4

Deep classifiers have achieved great success in visual recognition. However, real-world data is long-tailed by nature, leading to the mismatch between training and testing distribu…

cs.RO2018

3D Convolution on RGB-D Point Clouds for Accurate Model-free Object Pose Estimation

Zhongang Cai, Cunjun Yu, Quang-Cuong Pham

The conventional pose estimation of a 3D object usually requires the knowledge of the 3D model of the object. Even with the recent development in convolutional neural networks (CNN…

cs.RO2025

GSON: A Group-based Social Navigation Framework with Large Multimodal Model

Shangyi Luo, Peng Sun, Ji Zhu +4

With the increasing presence of service robots and autonomous vehicles in human environments, navigation systems need to evolve beyond simple destination reach to incorporate socia…

cs.RO2025

TeachingBot: Robot Teacher for Human Handwriting

Zhimin Hou, Cunjun Yu, David Hsu +1

Teaching and learning physical skills often require one-on-one interaction, making it difficult to scale up, as there are not enough human teachers. Robots offer an attractive alte…

cs.CV2020

Leveraging Localization for Multi-camera Association

Zhongang Cai, Cunjun Yu, Junzhe Zhang +2

We present McAssoc, a deep learning approach to the as-sociation of detection bounding boxes in different views ofa multi-camera system. The vast majority of the academiahas been d…

cs.LG2020

Balanced Activation for Long-tailed Visual Recognition

Jiawei Ren, Cunjun Yu, Zhongang Cai +1

Deep classifiers have achieved great success in visual recognition. However, real-world data is long-tailed by nature, leading to the mismatch between training and testing distribu…

cs.CV2020

Spatio-Temporal Graph Transformer Networks for Pedestrian Trajectory Prediction

Cunjun Yu, Xiao Ma, Jiawei Ren +2

Understanding crowd motion dynamics is critical to real-world applications, e.g., surveillance systems and autonomous driving. This is challenging because it requires effectively m…

cs.RO2019

Siamese Convolutional Neural Network for Sub-millimeter-accurate Camera Pose Estimation and Visual Servoing

Cunjun Yu, Zhongang Cai, Hung Pham +1

Visual Servoing (VS), where images taken from a camera typically attached to the robot end-effector are used to guide the robot motions, is an important technique to tackle robotic…