3 papers
cs.RO2025
NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?
Jierui Peng, Yanyan Zhang, Yicheng Duan +3
The evaluation of Vision-Language-Action (VLA) agents is hindered by the coarse, end-task success metric that fails to provide precise skill diagnosis or measure robustness to real…
cs.RO2025
A Navigation Framework Utilizing Vision-Language Models
Yicheng Duan, Kaiyu tang
Vision-and-Language Navigation (VLN) presents a complex challenge in embodied AI, requiring agents to interpret natural language instructions and navigate through visually rich, un…
cs.CV2025
Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)
Yicheng Duan, Xi Huang, Duo Chen
The rapid growth of video content demands efficient and precise retrieval systems. While vision-language models (VLMs) excel in representation learning, they often struggle with ad…