3 papers
cs.RO2026
Gripper-aware Vision Language Action Models
Hanyi Zhang, Zihong Luo, Tianyu Li +16
Vision language action models (VLAs) have advanced general purpose robotic grasping and manipulation by enabling robots to interpret visual observations and natural language instru…
cs.CV2026
RDVSv2: A Large-scale Benchmark for RGB-D Video Salient Object Detection
Tianyu Li, Jiahao He, Keren Fu +1
We introduce RDVSv2, a large-scale benchmark for RGB-D video salient object detection (RGB-D VSOD) with dense frame-level annotations. Existing datasets in this emerging field are…
cs.RO2026
Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution
Hanyi Zhang, Khang Nguyen, Charith Munasinghe +10
Robust robotic grasping remains a fundamental challenge for complex real-world applications. Recent advances in large-scale models demonstrate promising capabilities for reasoning…