14 papers
Twins: Learn to Predict Unified Representations with Focal Loss
Kaixiong Gong, Xin Cai, Bin Lin +9
Unified multimodal models seek a shared visual token space that supports both multimodal understanding and image generation. Discrete methods unify the interface via a shared codeb…
ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning
Xuanhua He, Jiaxin Xie, Mingzhe Zheng +1
Monocular video depth estimation requires temporal consistency, geometric accuracy, and generalization across diverse scenarios, yet existing methods struggle to achieve all three…
D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
Dengyang Jiang, Xin Jin, Dongyang Liu +9
The landscape of high-performance image generation models is currently shifting from the inefficient multi-step ones to the efficient few-step counterparts (e.g, Z-Image-Turbo and…
FastVMT: Eliminating Redundancy in Video Motion Transfer
Yue Ma, Zhikai Wang, Tianhao Ren +9
Video motion transfer aims to synthesize videos by generating visual content according to a text prompt while transferring the motion pattern observed in a reference video. Recent…
Group Editing: Edit Multiple Images in One Go
Yue Ma, Xinyu Wang, Qianli Ma +9
In this paper, we tackle the problem of performing consistent and unified modifications across a set of related images. This task is particularly challenging because these images m…
Manifold-Aware Exploration for Reinforcement Learning in Video Generation
Mingzhe Zheng, Weijie Kong, Yue Wu +9
Group Relative Policy Optimization (GRPO) methods for video generation like FlowGRPO remain far less reliable than their counterparts for language models and images. This gap arise…