From the 1 of 16 linked papers with an AI index.
16 papers
ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes
Xinrui Lin, Sha Zhang, Shumin Wang +3
Task-driven 3D affordance grounding aims to localize the functional region in a cluttered 3D scene that enables an action specified by a natural-language instruction. Existing meth…
ReTouch: Empowering Contact-Rich Dexterous Manipulation with Online-Refined Tactile Prediction
Shiqi Zhang, Xin Zhang, Yedong Shen +9
Fusing tactile signals has proven effective for contact-rich manipulation, enabling robots to perceive contact states and adapt to rapidly changing physical interactions. Yet effec…
ViCo3D: Empowering LiDAR-based Collaborative 3D Object Detection with Vision Foundation Models
Haojie Ren, Songrui Luo, Lingfeng Wang +6
The paper introduces ViCo3D, a framework that leverages vision foundation models to enrich LiDAR bird's-eye-view features for collaborative 3D object detection in V2X scenarios, ac…
GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation
Yuan Zhang, Shiqi Zhang, Yedong Shen +11
Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot emb…
Drift-Based Policy Optimization: Native One-Step Policy Learning for Online Robot Control
Yuxuan Gao, Yedong Shen, Shiqi Zhang +6
Although multi-step generative policies achieve strong performance in robotic manipulation by modeling multimodal action distributions, they require multi-step iterative denoising…
GA-GS: Generation-Assisted Gaussian Splatting for Static Scene Reconstruction
Yedong Shen, Shiqi Zhang, Sha Zhang +6
Reconstructing static 3D scene from monocular video with dynamic objects is important for numerous applications such as virtual reality and autonomous driving. Current approaches t…