From the 1 of 12 linked papers with an AI index.
12 papers
Native Video-Action Pretraining for Generalizable Robot Control
Qihang Zhang, Lin Li, Luyao Zhang +26
The paper introduces LingBot-VA 2.0, a video-action foundation model designed specifically for robot control, featuring a semantic visual-action tokenizer, causal pretraining, a sp…
A Closed-Loop Multi-Agent Framework for Robust Multi-Robot Manipulation
Yi-Xiang He, Lan Wei, Haoming Cen +6
Multi-robot systems provide the parallelism and redundancy necessary for long-horizon tasks, while Large Language Models (LLMs) offer the reasoning capabilities to decompose these…
RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation
Chuanrui Zhang, Zhengxian Wu, Guanxing Lu +2
Learned world models hold significant potential as neural simulators for robotic manipulation. However, prevalent 2D video-based models inherently lack the spatial and kinematic re…
BrickCraft: Visuomotor Skill Composition with Situated Manual Guidance for Long-Horizon Interlocking Brick Assembly
Jichuan Yu, Bowei Li, Zhenran Tang +4
Autonomous robotic assembly of interlocking bricks demands seamless integration of long-horizon task reasoning, spatial grounding, and fine-grained manipulation. This paper present…
RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation
Yuquan Xue, Guanxing Lu, Zhenyu Wu +4
Vision-Language-Action (VLA) models have shown strong manipulation capability when trained with large-scale imitation learning datasets. However, these datasets that predominantly…
VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
Wenkai Guo, Guanxing Lu, Haoyuan Deng +3
Vision-Language-Action models (VLAs) achieve strong performance in general robotic manipulation tasks by scaling imitation learning. However, existing VLAs are limited to predictin…