11 papers
Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?
Guoheng Sun, Kaixi Feng, Shwai He +8
Vision-Language-Action (VLA) models enable instruction-driven robotic manipulation, but they inherit oversized language backbones from pretrained VLMs whose capacity far exceeds wh…
VLA Knows Its Limits: Adaptive Execution Horizons for Robot Policies
Haoxuan Wang, Gengyu Zhang, Yan Yan +2
Action chunking has recently emerged as a standard practice in flow-based Vision-Language-Action (VLA) models. However, the effect and choice of the execution horizon - the number…
When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making
Jun Liu, Pu Zhao, Zhenglun Kong +12
Embodied robotic systems increasingly rely on large language model (LLM)-based agents to support high-level reasoning, planning, and decision-making during interactions with the en…
Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems
Jiabao Ji, Yongchao Chen, Yang Zhang +4
Multi-robot control in cluttered environments is a challenging problem that involves complex physical constraints, including robot-robot collisions, robot-obstacle collisions, and…
Attention-aware Inference Optimizations for Large Vision-Language Models with Memory-efficient Decoding
Fatih Ilhan, Gaowen Liu, Ramana Rao Kompella +5
Large Vision-Language Models (VLMs) have achieved remarkable success in multi-modal reasoning, but their inference time efficiency remains a significant challenge due to the memory…
A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoning
Yichang Xu, Gaowen Liu, Ramana Rao Kompella +6
This paper presents a multi-agent perception-action exploration alliance, dubbed A4VL, for efficient long-video reasoning. A4VL operates in a multi-round perception-action explorat…