4 papers
AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models
Xiaoqi Li, Muhe Cai, Jiadong Xu +5
Vision-Language-Action (VLA) models have significantly advanced the capabilities of robotic agents in executing diverse tasks; however, they still face challenges in contact-rich m…
SR3D: Unleashing Single-view 3D Reconstruction for Transparent and Specular Object Grasping
Mingxu Zhang, Xiaoqi Li, Jiahui Xu +5
Recent advancements in 3D robotic manipulation have improved grasping of everyday objects, but transparent and specular materials remain challenging due to depth sensing limitation…
SpatialBot: Precise Spatial Understanding with Vision Language Models
Wenxiao Cai, Iaroslav Ponomarenko, Jianhao Yuan +4
Vision Language Models (VLMs) have achieved impressive performance in 2D image understanding, however they are still struggling with spatial understanding which is the foundation o…
TransDiff: Diffusion-Based Method for Manipulating Transparent Objects Using a Single RGB-D Image
Haoxiao Wang, Kaichen Zhou, Binrui Gu +6
Manipulating transparent objects presents significant challenges due to the complexities introduced by their reflection and refraction properties, which considerably hinder the acc…