collaborators

10 papers

cs.CV2026

Rosetta: Composable Native Multimodal Pretraining

Xiangyue Liu, Zijian Zhang, Miles Yang +3

Achieving true artificial general intelligence requires foundation models capable of integrating new modalities without forgetting prior knowledge. However, accommodating continuou…

cs.CV2026

Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding

Xiangyue Liu, Zijian Zhang, Miles Yang +3

Empowering Large Multimodal Models (LMMs) with image generation often leads to catastrophic forgetting in understanding tasks due to severe gradient conflicts. While existing parad…

cs.CV2026

MeshFlow: Efficient Artistic Mesh Generation via MeshVAE and Flow-based Diffusion Transformer

Weiyu Li, Antoine Toisoul, Tom Monnier +4

We present MeshFlow, a new method for generating artist-like 3D meshes. Current mesh generators often adopt Auto-Regressive (AR) next-token prediction, a natural choice given the d…

cs.CV2026

AutoWeather4D: Autonomous Driving Video Weather Conversion via G-Buffer Dual-Pass Editing

Tianyu Liu, Weitao Xiong, Kunming Luo +4

Generative video models have significantly advanced the photorealistic synthesis of adverse weather for autonomous driving; however, they consistently demand massive datasets to le…

cs.CV2026

PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation

Hongyu Yan, Kunming Luo, Weiyu Li +5

Pose stylization, which aims to synthesize stylized content aligning with target poses, serves as a fundamental task across 2D, 3D, and video domains. In the 3D realm, prevailing a…

cs.CV2025

SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model

Yukai Shi, Weiyu Li, Zihao Wang +4

We propose a decoupled 3D scene generation framework called SceneMaker in this work. Due to the lack of sufficient open-set de-occlusion and pose estimation priors, existing method…