activity
20242026
collaborators

10 papers

cs.CV2026

ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes

Xinrui Lin, Sha Zhang, Shumin Wang +3

Task-driven 3D affordance grounding aims to localize the functional region in a cluttered 3D scene that enables an action specified by a natural-language instruction. Existing meth…

cs.RO2026

ReTouch: Empowering Contact-Rich Dexterous Manipulation with Online-Refined Tactile Prediction

Shiqi Zhang, Xin Zhang, Yedong Shen +9

Fusing tactile signals has proven effective for contact-rich manipulation, enabling robots to perceive contact states and adapt to rapidly changing physical interactions. Yet effec…

cs.CV2026

GA-GS: Generation-Assisted Gaussian Splatting for Static Scene Reconstruction

Yedong Shen, Shiqi Zhang, Sha Zhang +6

Reconstructing static 3D scene from monocular video with dynamic objects is important for numerous applications such as virtual reality and autonomous driving. Current approaches t…

cs.RO2025

LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents

Rui Li, Zixuan Hu, Wenxi Qu +13

Scientific embodied agents play a crucial role in modern laboratories by automating complex experimental workflows. Compared to typical household environments, laboratory settings…

cs.AI2025

VLMPlanner: Integrating Visual Language Models with Motion Planning

Zhipeng Tang, Sha Zhang, Jiajun Deng +5

Integrating large language models (LLMs) into autonomous driving motion planning has recently emerged as a promising direction, offering enhanced interpretability, better controlla…

cs.CV2025

PGOV3D: Open-Vocabulary 3D Semantic Segmentation with Partial-to-Global Curriculum

Shiqi Zhang, Sha Zhang, Jiajun Deng +3

Existing open-vocabulary 3D semantic segmentation methods typically supervise 3D segmentation models by merging text-aligned features (e.g., CLIP) extracted from multi-view images…