activity
20242026
collaborators

7 papers

cs.RO2026

PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution

Yang Liu, Weixing Chen, Xinshuai Song +8

Vision-language-action models, world models, and agentic planners each advance physical intelligence, yet their composition lacks a common execution abstraction, shared state, sema…

cs.RO2026

Bridge-WA: Predicting Where and How the World Changes for Robotic Action

Yongjie Bai, Hanting Wang, Mingtong Dai +3

General-purpose vision-language-action models benefit from large vision-language priors, but effective manipulation also requires anticipating action-relevant scene changes. Existi…

cs.RO2026

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models

Jingzhou Luo, Yifan Wen, Yongjie Bai +3

Vision-Language-Action (VLA) models have shown strong performance on embodied manipulation, yet they remain brittle under visual observation changes, paraphrased language instructi…

cs.RO2026

SkiP: When to Skip and When to Refine for Efficient Robot Manipulation

Mingtong Dai, Guanqi Peng, Yongjie Bai +5

Previous imitation learning policies predict future actions at every control step, whether in smooth motion phases or precise, contact-rich operation phases. This uniform treatment…

cs.RO2025

RoVer: Robot Reward Model as Test-Time Verifier for Vision-Language-Action Model

Mingtong Dai, Lingbo Liu, Yongjie Bai +6

Vision-Language-Action (VLA) models have become a prominent paradigm for embodied intelligence, yet further performance improvements typically rely on scaling up training data and…

cs.RO2025

Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation

Yongjie Bai, Zhouxia Wang, Yang Liu +8

Recent vision-language-action (VLA) models for multi-task robot manipulation often rely on fixed camera setups and shared visual encoders, which limit their performance under occlu…