collaborators

20 papers

cs.RO2026

PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution

Yang Liu, Weixing Chen, Xinshuai Song +8

Vision-language-action models, world models, and agentic planners each advance physical intelligence, yet their composition lacks a common execution abstraction, shared state, sema…

cs.RO2026

In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics

Xiaomeng Fu, Junfan Lin, Yang Liu +4

Synthesizing human motion from textual descriptions is essential for immersive digital applications, yet existing methods face a persistent trade-off between semantic fidelity and…

cs.CV2026

EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence

Linpeng Huang, Weixing Chen, Zexin Chen +2

Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA). Nevertheless, existing benchmarks are predomin…

cs.CV2026

When Preference Labels Fall Short: Aligning Diffusion Models from Real Data

Weiyan Chen, Weijian Deng, Yao Xiao +5

Preference alignment aims to guide generative models by learning from comparisons between preferred and non-preferred samples. In practice, most existing approaches rely on prefere…

cs.CV2026

PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation

Weixing Chen, Zhuoqian Feng, Yang Liu +6

Generating physically consistent 3D tabletop scenes is a fundamental yet underexplored problem for interactive and generalist robotic learning. The challenge stems from dense objec…

cs.AI2026

PolarMem: A Training-Free Polarized Latent Graph Memory for Verifiable Vision-Language Models

Zhisheng Chen, Tingyu Wu, Zijie Zhou +7

Memory is not merely a storage mechanism for intelligent systems, but a structure for organizing evidence and constraining belief. This is especially important for multimodal reaso…