5 papers
CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion
Dianbing Xi, Jiepeng Wang, Yuanzhi Liang +8
We tackle the dual challenges of video understanding and controllable video generation within a unified diffusion framework. Our key insights are two-fold: geometry-only cues (e.g.…
PFAvatar: Pose-Fusion 3D Personalized Avatar Reconstruction from Real-World Outfit-of-the-Day Photos
Dianbing Xi, Guoyuan An, Jingsen Zhu +6
We propose PFAvatar (Pose-Fusion Avatar), a new method that reconstructs high-quality 3D avatars from Outfit of the Day(OOTD) photos, which exhibit diverse poses, occlusions, and c…
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
Dianbing Xi, Jiepeng Wang, Yuanzhi Liang +5
In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff , aiming to synthesize and comprehend multiple video visual content in a single diffusion mo…
Inverse Rendering using Multi-Bounce Path Tracing and Reservoir Sampling
Yuxin Dai, Qi Wang, Jingsen Zhu +4
We present MIRReS, a novel two-stage inverse rendering framework that jointly reconstructs and optimizes the explicit geometry, material, and lighting from multi-view images. Unlik…
SGW-based Multi-Task Learning in Vision Tasks
Ruiyuan Zhang, Yuyao Chen, Yuchi Huo +4
Multi-task-learning(MTL) is a multi-target optimization task. Neural networks try to realize each target using a shared interpretative space within MTL. However, as the scale of da…