5 papers
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
Juyi Lin, Amir Taherin, Arash Akbari +11
Recent large-scale Vision Language Action (VLA) models have shown superior performance in robotic manipulation tasks guided by natural language. However, current VLA models suffer…
Next Best View Selections for Semantic and Dynamic 3D Gaussian Splatting
Yiqian Li, Wen Jiang, Kostas Daniilidis
Understanding semantics and dynamics has been crucial for embodied agents in various tasks. Both tasks have much more data redundancy than the static scene understanding task. We f…
PEVLM: Parallel Encoding for Vision-Language Models
Letian Kang, Shixian Luo, Yiqiang Li +5
Vision-Language Models (VLMs) have demonstrated strong capabilities in multimodal understanding and generation tasks. However, their application to long video understanding remains…
Moral Reasoning Across Languages: The Critical Role of Low-Resource Languages in LLMs
Huichi Zhou, Zehao Xu, Munan Zhao +3
In this paper, we introduce the Multilingual Moral Reasoning Benchmark (MMRB) to evaluate the moral reasoning abilities of large language models (LLMs) across five typologically di…
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
Dongping Chen, Yue Huang, Siyuan Wu +17
Recently, Multimodal Large Language Models (MLLMs) have been used as agents to control keyboard and mouse inputs by directly perceiving the Graphical User Interface (GUI) and gener…