collaborators

5 papers

cs.RO2026

Data Pyramid for Embodied Manipulation: A Survey

Yifan Ye, Yankai Fu, Yaoxu Lv +26

Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations w…

cs.CV2026

Efficient-VLN: A Simple yet Strong Baseline for Efficient Vision-Language Navigation

Duo Zheng, Shijia Huang, Yanyang Li +1

While Multimodal Large Language Models (MLLMs) have demonstrated significant promise in Vision-Language Navigation (VLN), existing agents remain heavily constrained by systemic bot…

cs.CV2025

Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors

Duo Zheng, Shijia Huang, Yanyang Li +1

Previous research has investigated the application of Multimodal Large Language Models (MLLMs) in understanding 3D scenes by interpreting them as videos. These approaches generally…

cs.CL2025

CLEVA: Toward Comprehensive and Contamination-Free Language Model Evaluation

Yanyang Li, Tin Long Wong, Cheung To Hung +5

Recent advances in large language models (LLMs) have shown significant promise, yet their evaluation raises concerns, particularly regarding data contamination due to the lack of a…

cs.CV2025

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Duo Zheng, Shijia Huang, Liwei Wang

The rapid advancement of Multimodal Large Language Models (MLLMs) has significantly impacted various multimodal tasks. However, these models face challenges in tasks that require s…