3 papers
cs.CV2026
Teaching Vision-Language-Action Models What to See and Where to Look
Yuguang Yang, Canyu Chen, Zhewen Tan +10
Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing VLAs' training relies heavily on text-centric visual q…
cs.CV2025
Semore: VLM-guided Enhanced Semantic Motion Representations for Visual Reinforcement Learning
Wentao Wang, Chunyang Liu, Kehua Sheng +2
The growing exploration of Large Language Models (LLM) and Vision-Language Models (VLM) has opened avenues for enhancing the effectiveness of reinforcement learning (RL). However,…
cs.AI2025
Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph
Wentao Wang, Heqing Zou, Tianze Luo +8
Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated strong semantic understanding capabilities, but struggles to perform precise spatio-temporal understand…