4 papers
RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation
Dayu Xia, Yue Shi, Yao Mu +7
Vision-language models (VLMs) are increasingly explored as visual critics, reward generators, and failure detectors in robotic manipulation. These roles implicitly require models t…
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
Yang Liu, Pengxiang Ding, Tengyue Jiang +10
Vision-Language-Action (VLA) models map visual observations and natural-language instructions to robot actions; however, hierarchical and autoregressive paradigms often incur archi…
AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning
Dejie Yang, Zijing Zhao, Yang Liu
Visual Robot Manipulation (VRM) aims to enable a robot to follow natural language instructions based on robot states and visual observations, and therefore requires costly multi-mo…
PlanLLM: Video Procedure Planning with Refinable Large Language Models
Dejie Yang, Zijing Zhao, Yang Liu
Video procedure planning, i.e., planning a sequence of action steps given the video frames of start and goal states, is an essential ability for embodied AI. Recent works utilize L…