2 papers
cs.CV2026
DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing
Zhaokai Wang, Mingxin Liu, Zirun Zhu +11
Recent image generation and editing models can produce visually appealing natural images, yet they remain unreliable when the target image is a knowledge-intensive diagram whose co…
cs.GR2025
X2Video: Adapting Diffusion Models for Multimodal Controllable Neural Video Rendering
Zhitong Huang, Mohan Zhang, Renhan Wang +3
We present X2Video, the first diffusion model for rendering photorealistic videos guided by intrinsic channels including albedo, normal, roughness, metallicity, and irradiance, whi…