5 papers
Efficient-VLN: A Simple yet Strong Baseline for Efficient Vision-Language Navigation
Duo Zheng, Shijia Huang, Yanyang Li +1
While Multimodal Large Language Models (MLLMs) have demonstrated significant promise in Vision-Language Navigation (VLN), existing agents remain heavily constrained by systemic bot…
Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
Duo Zheng, Shijia Huang, Yanyang Li +1
Previous research has investigated the application of Multimodal Large Language Models (MLLMs) in understanding 3D scenes by interpreting them as videos. These approaches generally…
Learning to Reason from Feedback at Test-Time
Yanyang Li, Michael Lyu, Liwei Wang
Solving complex tasks in a single attempt is challenging for large language models (LLMs). Iterative interaction with the environment and feedback is often required to achieve succ…
CLEVA: Toward Comprehensive and Contamination-Free Language Model Evaluation
Yanyang Li, Tin Long Wong, Cheung To Hung +5
Recent advances in large language models (LLMs) have shown significant promise, yet their evaluation raises concerns, particularly regarding data contamination due to the lack of a…
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Duo Zheng, Shijia Huang, Liwei Wang
The rapid advancement of Multimodal Large Language Models (MLLMs) has significantly impacted various multimodal tasks. However, these models face challenges in tasks that require s…