works on

From the 1 of 16 linked papers with an AI index.

collaborators

16 papers

cs.CV2026

ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes

Xinrui Lin, Sha Zhang, Shumin Wang +3

Task-driven 3D affordance grounding aims to localize the functional region in a cluttered 3D scene that enables an action specified by a natural-language instruction. Existing meth…

cs.RO2026

ReTouch: Empowering Contact-Rich Dexterous Manipulation with Online-Refined Tactile Prediction

Shiqi Zhang, Xin Zhang, Yedong Shen +9

Fusing tactile signals has proven effective for contact-rich manipulation, enabling robots to perceive contact states and adapt to rapidly changing physical interactions. Yet effec…

cs.CV2026

ViCo3D: Empowering LiDAR-based Collaborative 3D Object Detection with Vision Foundation Models

Haojie Ren, Songrui Luo, Lingfeng Wang +6

The paper introduces ViCo3D, a framework that leverages vision foundation models to enrich LiDAR bird's-eye-view features for collaborative 3D object detection in V2X scenarios, ac…

cs.RO2026

GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation

Yuan Zhang, Shiqi Zhang, Yedong Shen +11

Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot emb…

cs.RO2026

Drift-Based Policy Optimization: Native One-Step Policy Learning for Online Robot Control

Yuxuan Gao, Yedong Shen, Shiqi Zhang +6

Although multi-step generative policies achieve strong performance in robotic manipulation by modeling multimodal action distributions, they require multi-step iterative denoising…

cs.CV2026

GA-GS: Generation-Assisted Gaussian Splatting for Static Scene Reconstruction

Yedong Shen, Shiqi Zhang, Sha Zhang +6

Reconstructing static 3D scene from monocular video with dynamic objects is important for numerous applications such as virtual reality and autonomous driving. Current approaches t…