works on

From the 1 of 13 linked papers with an AI index.

collaborators

13 papers

cs.CV2026

PhiZero: A World Model Built Around Physical Language

Shuyao Shang, Yuqi Wang, Ruopeng Gao +4

PhiZero is a physical world model that learns a compact discrete "physical language" from videos to predict future world states as language sequences before rendering them into rea…

cs.CV2026

SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion

Ruoyu Feng, Jinming Liu, Yuqi Wang +7

Training image generation foundation models consumes substantial resources. Previous methods have attempted to leverage semantic guidance to accelerate the training process, yet th…

cs.CV2026

Bridging Video Understanding and Generation in a Unified Framework

Yuqi Wang, Runyi Li, Ruoyu Feng +3

Recently, unified image generation and understanding have been extensively explored. However, extending such unified modeling paradigms to the video domain remains largely underexp…

cs.CV2026

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining

Juncheng Ma, Jianxin Bi, Yufan Deng +19

Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real-robot trajectories remai…

cs.CV2026

RL: Reasoning 3D Layouts from Relative Spatial Relations

Zhifeng Gu, Yuqi Wang, Bing Wang

Relative spatial relations provide a compact representation of spatial structure and are fundamental to relative spatial reasoning in 3D layout generation. Recent works leverage Mu…

cs.CV2026

Generation Navigator: A State-Aware Agentic Framework for Image Generation

Jinming Liu, Ruoyu Feng, Yuqi Wang +2

Despite rapid advances in text-to-image generation, faithfully realizing user intent remains challenging, often requiring manual multi-turn trial and error. To automate this proces…