6 papers
TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model
Yabo Chen, Yuanzhi Liang, Jiepeng Wang +24
World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent vi…
CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion
Dianbing Xi, Jiepeng Wang, Yuanzhi Liang +8
We tackle the dual challenges of video understanding and controllable video generation within a unified diffusion framework. Our key insights are two-fold: geometry-only cues (e.g.…
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
Chi Zhang, Jiepeng Wang, Youming Wang +5
We present UniModel, a unified generative model that jointly supports visual understanding and visual generation within a single pixel-to-pixel diffusion framework. Our goal is to…
Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction
Chenyou Fan, Fangzheng Yan, Chenjia Bai +4
Learning a generalizable bimanual manipulation policy is extremely challenging for embodied agents due to the large action space and the need for coordinated arm movements. Existin…
Metric-Solver: Sliding Anchored Metric Depth Estimation from a Single Image
Tao Wen, Jiepeng Wang, Yabo Chen +3
Accurate and generalizable metric depth estimation is crucial for various computer vision applications but remains challenging due to the diverse depth scales encountered in indoor…
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
Dianbing Xi, Jiepeng Wang, Yuanzhi Liang +5
In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff , aiming to synthesize and comprehend multiple video visual content in a single diffusion mo…