7 papers
SSI-Policy: Learning Structured Scene Interfaces for Vision-Language Robotic Manipulation
Kaijun Wang, Zikai Ouyang, Xuping Wu +6
Real-world robotic manipulation demands spatial grounding, task-aware reasoning, and precise control. Learning such capabilities becomes particularly challenging in the low-data re…
AISPO: Enhancing Depth Reliability for Robotic Manipulation of Non-Lambertian Objects via Affine-Invariant Shape Prior
Zhiming Chen, Linfang Zheng, Kun Zhang +4
Reliable depth perception is critical for robotic manipulation, especially for non-Lambertian objects such as transparent or highly specular surfaces, where raw depth measurements…
From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data
Linfang Zheng, Zikai Ouyang, Chen Wang +2
Video is a scalable observation of physical dynamics: it captures how objects move, how contact unfolds, and how scenes evolve under interaction -- all without requiring robot acti…
Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation
Chuye Zhang, Xiaoxiong Zhang, Wei Pan +2
Robotic manipulation in unstructured environments requires systems that can generalize across diverse tasks while maintaining robust and reliable performance. We introduce {GVF-TAP…
Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics
Tze Ho Elden Tse, Runyang Feng, Linfang Zheng +5
With the availability of egocentric 3D hand-object interaction datasets, there is increasing interest in developing unified models for hand-object pose estimation and action recogn…
Object Gaussian for Monocular 6D Pose Estimation from Sparse Views
Luqing Luo, Shichu Sun, Jiangang Yang +3
Monocular object pose estimation, as a pivotal task in computer vision and robotics, heavily depends on accurate 2D-3D correspondences, which often demand costly CAD models that ma…