activity
20242026
collaborators

17 papers

cs.RO2026

PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution

Yang Liu, Weixing Chen, Xinshuai Song +8

Vision-language-action models, world models, and agentic planners each advance physical intelligence, yet their composition lacks a common execution abstraction, shared state, sema…

cs.CV2026

EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence

Linpeng Huang, Weixing Chen, Zexin Chen +2

Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA). Nevertheless, existing benchmarks are predomin…

cs.CV2026

PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation

Weixing Chen, Zhuoqian Feng, Yang Liu +6

Generating physically consistent 3D tabletop scenes is a fundamental yet underexplored problem for interactive and generalist robotic learning. The challenge stems from dense objec…

cs.CV2026

LASAR: Towards Spatio-temporal Reasoning with Latent Cognitive Map

Jinzhou Tang, Sidi Liu, Waikit Xiu +2

A fundamental challenge in embodied AI is verifying if agents build internal models of spatial structure or merely learn to mimic task-specific expert trajectories. This is critica…

cs.RO2026

Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation

Yongjie Bai, Zhouxia Wang, Yang Liu +8

Recent vision-language-action (VLA) models for multi-task robot manipulation often rely on fixed camera setups and shared visual encoders, which limit their performance under occlu…

cs.CV2026

DDP-WM: Disentangled Dynamics Prediction for Efficient World Models

Shicheng Yin, Kaixuan Yin, Weixing Chen +3

World models are essential for autonomous robotic planning. However, the substantial computational overhead of existing dense Transformerbased models significantly hinders real-tim…