9 papers
LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding
Zhewei Zhang, Puyue Wang, Guanren Qiao +10
Vision-Language-Action (VLA) models transform representations from pretrained vision-language models (VLMs) into robot actions, yet the interface that routes intermediate VLM featu…
Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?
Jingtao He, Hongliang Lu, Xiaoyun Qiu +2
Vision-Language-Action (VLA) models have demonstrated promising capability in autonomous driving, highlighting the potential of unified multimodal architectures for jointly modelin…
UniT: Unified Geometry Learning with Group Autoregressive Transformer
Haotian Wang, Yusong Huang, Zhaonian Kuang +4
Recent feed-forward models have significantly advanced geometry perception for inferring dense 3D structure from sensor observations. However, its essential capabilities remain fra…
Accelerating Rectified Flow Models via Trajectory-Aware Caching
Xiao Liu, Kai Liu, Naiyang Guan +5
Diffusion and rectified flow (RF) models generate high-fidelity images and videos, but their iterative velocity-field evaluations are computationally expensive. Existing caching me…
The Great March 100: 100 Detail-oriented Tasks for Evaluating Embodied AI Agents
Ziyu Wang, Chenyuan Liu, Yushun Xiang +16
Recently, with the rapid development of robot learning and imitation learning, numerous datasets and methods have emerged. However, these datasets and their task designs often lack…
Coordinated Pandemic Control with Large Language Model Agents as Policymaking Assistants
Ziyi Shi, Xusen Guo, Hongliang Lu +7
Effective pandemic control requires timely and coordinated policymaking across administrative regions that are intrinsically interdependent. However, human-driven responses are oft…