collaborators

10 papers

cs.AI2026

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

Ying Chen, Weizhen Li, Zhe Hu +7

Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state…

cs.AI2026

Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction

Jialian Li, Junhong Liu, Yuchen Cao +6

Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge. As embodied agents become increasingly capable, th…

cs.RO2026

Robots as Tokens: Unified Diffusion Transformer for Coordinated Multi-Robot Trajectory Generation

Ruofei Bai, Jie Chen, Yuxin Cai +3

The success of generative models in language and visual generation has inspired extensive applications to generative robot planning. However, most existing works either focus on si…

cs.AI2026

Engagement Process: Rethinking the Temporal Interface of Action and Observation

Jialian Li, Yuchen Cao, Junhong Liu +5

Task completion in digital and physical environments increasingly involves complex temporal interaction, where actions and observations unfold over different time scales rather tha…

cs.AI2026

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents

Ying Chen, Lihuang Fang, Rui Jiang +4

Standard embodied evaluations do not independently score whether an agent correctly commits to task completion at episode closure, a capacity we call terminal commitment. Behaviora…

cs.RO2026

ImagiNav: Scalable Embodied Navigation via Generative Visual Prediction and Inverse Dynamics

Jie Chen, Yuxin Cai, Yizhuo Wang +5

Enabling robots to navigate open-world environments via natural language is critical for general-purpose autonomy. Yet, Vision-Language Navigation has relied on end-to-end policies…