16 papers
RoboStereo: Dual-Tower 4D Embodied World Models for Unified Policy Optimization
Ruicheng Zhang, Guangyu Chen, Zunnan Xu +5
Scalable Embodied AI faces fundamental constraints due to prohibitive costs and safety risks of real-world interaction. While Embodied World Models (EWMs) offer promise through ima…
Identity-Consistent Video Generation under Large Facial-Angle Variations
Bin Hu, Zipeng Qi, Guoxi Huang +6
Single-view reference-to-video methods often struggle to preserve identity consistency under large facial-angle variations. This limitation naturally motivates the incorporation of…
Controllable Layer Decomposition for Reversible Multi-Layer Image Generation
Zihao Liu, Zunnan Xu, Shi Shu +4
This work presents Controllable Layer Decomposition (CLD), a method for achieving fine-grained and controllable multi-layer separation of raster images. In practical workflows, des…
Igniting VLMs toward the Embodied Space
Andy Zhai, Brae Liu, Bruno Fang +17
While foundation models show remarkable progress in language and vision, existing vision-language models (VLMs) still have limited spatial and embodiment understanding. Transferrin…
Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time Training
Ruicheng Zhang, Jun Zhou, Zunnan Xu +5
Trajectory-Guided image-to-video (I2V) generation aims to synthesize videos that adhere to user-specified motion instructions. Existing methods typically rely on computationally ex…
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
Yicheng Xiao, Lin Song, Rui Yang +6
With the advancement of language models, unified multimodal understanding and generation have made significant strides, with model architectures evolving from separated components…