6 papers · 1 filter
PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution
Yang Liu, Weixing Chen, Xinshuai Song +8
Vision-language-action models, world models, and agentic planners each advance physical intelligence, yet their composition lacks a common execution abstraction, shared state, sema…
In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics
Xiaomeng Fu, Junfan Lin, Yang Liu +4
Synthesizing human motion from textual descriptions is essential for immersive digital applications, yet existing methods face a persistent trade-off between semantic fidelity and…
RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models
Jingzhou Luo, Yifan Wen, Yongjie Bai +3
Vision-Language-Action (VLA) models have shown strong performance on embodied manipulation, yet they remain brittle under visual observation changes, paraphrased language instructi…
Uncovering Linguistic Fragility in Vision-Language-Action Models via Diversity-Aware Red Teaming
Baoshun Tong, Haoran He, Ling Pan +2
Vision-Language-Action (VLA) models have achieved remarkable success in robotic manipulation. However, their robustness to linguistic nuances remains a critical, under-explored saf…
Referring-Aware Visuomotor Policy Learning for Closed-Loop Manipulation
Jiahua Ma, Yiran Qin, Xin Wen +5
This paper addresses a fundamental problem of visuomotor policy learning for robotic manipulation: how to enhance robustness in out-of-distribution execution errors or dynamically…
AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning
Weixing Chen, Dafeng Chi, Yang Liu +7
The automated generation of layouts is vital for embodied intelligence and autonomous systems, supporting applications from virtual environment construction to home robot deploymen…