2 papers
cs.RO2025
See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
Guangyan Chen, Meiling Wang, Qi Shao +10
Developing robust and general-purpose manipulation policies represents a fundamental objective in robotics research. While Vision-Language-Action (VLA) models have demonstrated pro…
cs.CV2024
Open-RGBT: Open-vocabulary RGB-T Zero-shot Semantic Segmentation in Open-world Environments
Meng Yu, Luojie Yang, Xunjie He +2
Semantic segmentation is a critical technique for effective scene understanding. Traditional RGB-T semantic segmentation models often struggle to generalize across diverse scenario…