7 papers
Turning Video Models into Generalist Robot Policies
Sizhe Lester Li, Evan Kim, Xingjian Bai +4
Video generative models have emerged as a promising robotics backbone, capable of generating videos that depict the completion of complex tasks across embodiments and environments.…
End-to-End Training for Unified Tokenization and Latent Denoising
Shivam Duggal, Xingjian Bai, Zongze Wu +5
Latent diffusion models (LDMs) enable high-fidelity synthesis by operating in learned latent spaces. However, training state-of-the-art LDMs requires complex staging: a tokenizer m…
Causality in Video Diffusers is Separable from Denoising
Xingjian Bai, Guande He, Zhengqi Li +3
Causality -- referring to temporal, uni-directional cause-effect relationships between components -- underlies many complex generative processes, including videos, language, and ro…
GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models
Jingxuan Wei, Caijun Jia, Xi Bai +7
The advent of Unified Multimodal Models (UMMs) signals a paradigm shift in artificial intelligence, moving from passive perception to active, cross-modal generation. Despite their…
Mean Flows for One-step Generative Modeling
Zhengyang Geng, Mingyang Deng, Xingjian Bai +2
We propose a principled and effective framework for one-step generative modeling. We introduce the notion of average velocity to characterize flow fields, in contrast to instantane…
Wasserstein distributional adversarial training for deep neural networks
Xingjian Bai, Guangyi He, Yifan Jiang +1
Design of adversarial attacks for deep neural networks, as well as methods of adversarial training against them, are subject of intense research. In this paper, we propose methods…