4 papers
Gripper-aware Vision Language Action Models
Hanyi Zhang, Zihong Luo, Tianyu Li +16
Vision language action models (VLAs) have advanced general purpose robotic grasping and manipulation by enabling robots to interpret visual observations and natural language instru…
Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution
Hanyi Zhang, Khang Nguyen, Charith Munasinghe +10
Robust robotic grasping remains a fundamental challenge for complex real-world applications. Recent advances in large-scale models demonstrate promising capabilities for reasoning…
PG-SAM: Prior-Guided SAM with Medical for Multi-organ Segmentation
Yiheng Zhong, Zihong Luo, Chengzhi Liu +7
Segment Anything Model (SAM) demonstrates powerful zero-shot capabilities; however, its accuracy and robustness significantly decrease when applied to medical image segmentation. E…
Incomplete Modality Disentangled Representation for Ophthalmic Disease Grading and Diagnosis
Chengzhi Liu, Zile Huang, Zhe Chen +6
Ophthalmologists typically require multimodal data sources to improve diagnostic accuracy in clinical decisions. However, due to medical device shortages, low-quality data and data…