Move-Then-Operate: Behavioral Phasing for Human-Like Robotic Manipulation
arXiv:2604.23620
The paper introduces Move-Then-Operate, a vision‑language framework that splits robotic manipulation into a coarse relocation phase and a contact‑critical interaction phase using a dual‑expert policy and a learnable phase selector, achieving higher success rates with less data and training time.
Abstract
We present Move-Then-Operate, a Vision language action framework that explicitly decouples robotic manipulation into two distinct behavioral phases: coarse relocation (move) and contact-critical interaction (operate). Unlike monolithic policies that conflate these heterogeneous regimes, our architecture employs a dual-expert policy routed by a learnable phase selector, introducing a structural inductive bias that isolates phase-specific dynamics. Phase labels are automatically generated via an MLLM-based pipeline conditioned on lightweight contextual cues such as end-effector velocity and subtask decomposition to ensure alignment with human motor patterns. Evaluated on the RoboTwin2 benchmark, our method achieves an average success rate of , outperforming the monolithic baseline by . It matches or exceeds models trained on more data and reaches peak performance in fewer training steps, demonstrating that architectural disentanglement of move and operate phases is a highly effective and efficient strategy for mastering high-precision manipulation.
15 pages, 10 figures