3 papers
cs.CV2025
A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
Chenxuan Li, Jiaming Liu, Guanqun Wang +8
Recently, some studies have integrated Multimodal Large Language Models into robotic manipulation, constructing vision-language-action models (VLAs) to interpret multimodal informa…
cs.CV2024
SCBench: A Sports Commentary Benchmark for Video LLMs
Kuangzhi Ge, Lingjun Chen, Kevin Zhang +6
Recently, significant advances have been made in Video Large Language Models (Video LLMs) in both academia and industry. However, methods to evaluate and benchmark the performance…
cs.CV2024
MR-MLLM: Mutual Reinforcement of Multimodal Comprehension and Vision Perception
Guanqun Wang, Xinyu Wei, Jiaming Liu +5
In recent years, multimodal large language models (MLLMs) have shown remarkable capabilities in tasks like visual question answering and common sense reasoning, while visual percep…