8 papers
OC-VLA++: Monocular Geometry-Guided Cross-View Consistency for Viewpoint-Robust Robotic Manipulation
Tianyi Zhang, Ziyang Gong, Zhenjie Yang +2
We propose OC-VLA++, an extension of OC-VLA for viewpoint generalization under limited camera coverage. While OC-VLA grounds robot actions in the camera coordinate system to align…
Hybrid LLM-based Intelligent Framework for Robot Task Scheduling
Swayamjit Saha, Subhabrata Das, Haonan Duan +1
This study introduces intelligent frameworks that use Large Language Models (LLMs) to improve task scheduling for construction robots. The LLM is fed with key data about the desire…
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
Yaolun Zhang, Ruohui Wang, Jiahao Wang +6
Video understanding with multimodal large language models (MLLMs) remains challenging due to the long token sequences of videos, which contain extensive temporal dependencies and r…
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
Zhi Hou, Tianyi Zhang, Yuwen Xiong +8
While recent vision-language-action models trained on diverse robot datasets exhibit promising generalization capabilities with limited in-domain data, their reliance on compact ac…
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
Tianyi Zhang, Haonan Duan, Haoran Hao +3
Vision-Language-Action (VLA) models frequently encounter challenges in generalizing to real-world environments due to inherent discrepancies between observation and action spaces.…
Meta-Designing Quantum Experiments with Language Models
Sören Arlt, Haonan Duan, Felix Li +3
Artificial Intelligence (AI) can solve complex scientific problems beyond human capabilities, but the resulting solutions offer little insight into the underlying physical principl…