3 papers
cs.RO2026
Advances and Innovations in the Multi-Agent Robotic System (MARS) Challenge
Li Kang, Heng Zhou, Xiufeng Song +41
Recent advancements in multimodal large language models and vision-languageaction models have significantly driven progress in Embodied AI. As the field transitions toward more com…
cs.CV2025
End-to-End Motion Capture from Rigid Body Markers with Geodesic Loss
Hai Lan, Zongyan Li, Jianmin Hu +2
Marker-based optical motion capture (MoCap), while long regarded as the gold standard for accuracy, faces practical challenges, such as time-consuming preparation and marker identi…
cs.CL2025
From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
Hala Sheta, Eric Huang, Shuyu Wu +9
We introduce VLM-Lens, a toolkit designed to enable systematic benchmarking, analysis, and interpretation of vision-language models (VLMs) by supporting the extraction of intermedi…