2 papers
cs.RO2026
HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving
Weizhe Tang, Junwei You, Jiaxi Liu +5
End-to-end autonomous driving models increasingly benefit from large vision--language models for semantic understanding, yet ensuring safe and accurate operation under long-tail co…
cs.CV2025
Toward Rich Video Human-Motion2D Generation
Ruihao Xi, Xuekuan Wang, Yongcheng Li +5
Generating realistic and controllable human motions, particularly those involving rich multi-character interactions, remains a significant challenge due to data scarcity and the co…