works on

From the 1 of 10 linked papers with an AI index.

collaborators

10 papers

cs.CV2026

MMOE: Modernizing Diffusion Transformers with Efficient Expert Design

Yanhao Jia, Jiepeng Wang, Haibin Huang +3

Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capacity grows. AIGC Foundation…

cs.CV2026

ShotPlan: Cinematic Video Generation with Learnable Planning Token

Su Guo, Guangce Liu, Haosen Yang +7

Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective mult…

cs.CV2026

SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models

Ruoyu Wang, Jialun Liu, Huayang Huang +5

The paper introduces Self-Imagination Fine-Tuning (SIFT), a method that trains video diffusion models on their own generated videos to improve physical realism and disentangle moti…

cs.CV2025

TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model

Yabo Chen, Yuanzhi Liang, Jiepeng Wang +24

World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent vi…

cs.CV2025

CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion

Dianbing Xi, Jiepeng Wang, Yuanzhi Liang +8

We tackle the dual challenges of video understanding and controllable video generation within a unified diffusion framework. Our key insights are two-fold: geometry-only cues (e.g.…

cs.CV2025

UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation

Chi Zhang, Jiepeng Wang, Youming Wang +5

We present UniModel, a unified generative model that jointly supports visual understanding and visual generation within a single pixel-to-pixel diffusion framework. Our goal is to…