6 papers
PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution
Yang Liu, Weixing Chen, Xinshuai Song +8
Vision-language-action models, world models, and agentic planners each advance physical intelligence, yet their composition lacks a common execution abstraction, shared state, sema…
Category-Adaptive Cross-Modal Semantic Refinement and Transfer for Open-Vocabulary Multi-Label Recognition
Haijing Liu, Tao Pu, Hefeng Wu +4
Benefiting from the generalization capability of CLIP, recent vision language pre-training (VLP) models have demonstrated the ability to capture a wide range of visual concepts in…
RADAR: Benchmarking Vision-Language-Action Generalization via Real-World Dynamics, Spatial-Physical Intelligence, and Autonomous Evaluation
Yuhao Chen, Zhihao Zhan, Xiaoxin Lin +11
VLA models have achieved remarkable progress in embodied intelligence; however, their evaluation remains largely confined to simulations or highly constrained real-world settings.…
Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal Intervention
Haijing Liu, Zhiyuan Song, Hefeng Wu +3
Egocentric Referring Video Object Segmentation (Ego-RVOS) aims to segment the specific object actively involved in a human action, as described by a language query, within first-pe…
Human-Centric Open-Future Task Discovery: Formulation, Benchmark, and Scalable Tree-Based Search
Zijian Song, Xiaoxin Lin, Tao Pu +3
Recent progress in robotics and embodied AI is largely driven by Large Multimodal Models (LMMs). However, a key challenge remains underexplored: how can we advance LMMs to discover…
DART: Dual Adaptive Refinement Transfer for Open-Vocabulary Multi-Label Recognition
Haijing Liu, Tao Pu, Hefeng Wu +2
Open-Vocabulary Multi-Label Recognition (OV-MLR) aims to identify multiple seen and unseen object categories within an image, requiring both precise intra-class localization to pin…