activity
20242026
collaborators

11 papers

cs.CV2026

Distill What RGB Can Recover: Privileged 3D Evidence for RGB-Only Vision-Language Models

Yanbin Hu, Jin Cui, Jun Ye +4

3D scene understanding requires reasoning about entity existence, spatial layout, and object relations, yet RGB images alone often provide insufficient 3D cues. Existing 3D-VLMs co…

cs.RO2026

AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis

Junjie Ye, Rong Xue, Basile Van Hoorick +4

The collection of large-scale and diverse robot demonstrations remains a major bottleneck for imitation learning, as real-world data acquisition is costly and simulators offer limi…

cs.RO2026

Duet: Dual-Robot Understanding via Efficient Teaching

Yiqi Zhao, Ruohai Ge, Celina Shiyu Wang +10

Dual-robot collaboration enables tasks that exceed the reach and payload of a single robot, such as collaboratively transporting objects across environments and executing coordinat…

cs.CV2026

VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation

Hongyang Du, Junjie Ye, Xiaoyan Cong +7

While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, often resulting in object deforma…

cs.RO2026

SIMPLE: Simulation-Based Policy Learning and Evaluation for Humanoid Loco-manipulation

Songlin Wei, Zhenhao Ni, Jie Liu +9

Humanoid foundation models are advancing faster than we can evaluate them. While real-world testing is expensive and difficult to reproduce, existing simulation benchmarks focus pr…

cs.RO2026

RoboDream: Compositional World Models for Scalable Robot Data Synthesis

Junjie Ye, Rong Xue, Basile Van Hoorick +6

Scaling robot learning requires large-scale, diverse demonstrations, yet real-world data collection via teleoperation remains prohibitively expensive and time-consuming. While vide…