6 papers · 1 filter
Reward Lightning: Fast Video Generation via Homologous Preference Distillation
Jiaxiang Cheng, Bing Ma, Xuhua Ren +4
Achieving simultaneous preference alignment and distillation acceleration in video diffusion models remains an open challenge. Existing methods optimize the two objectives over mis…
ProductWebGen: Benchmarking Multimodal Product Webpage Generation
Zhihong Liu, Siqi Kou, Zheng Li +5
Crafting a product display webpage from a source product image, along with layout and visual content instructions, holds significant practical value for domains such as marketing,…
HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
Qi Cai, Jingwen Chen, Chengmin Gao +22
The evolution of visual generative models has long been constrained by fragmented architectures relying on disjoint text encoders and external VAEs. In this report, we present HiDr…
Implicit Preference Alignment for Human Image Animation
Yuanzhi Wang, Xuhua Ren, Jiaxiang Cheng +5
Human image animation has witnessed significant advancements, yet generating high-fidelity hand motions remains a persistent challenge due to their high degrees of freedom and moti…
Phased One-Step Adversarial Equilibrium for Video Diffusion Models
Jiaxiang Cheng, Bing Ma, Xuhua Ren +7
Video diffusion generation suffers from critical sampling efficiency bottlenecks, particularly for large-scale models and long contexts. Existing video acceleration methods, adapte…
PersonaVlog: Personalized Multimodal Vlog Generation with Multi-Agent Collaboration and Iterative Self-Correction
Xiaolu Hou, Bing Ma, Jiaxiang Cheng +5
With the growing demand for short videos and personalized content, automated Video Log (Vlog) generation has become a key direction in multimodal content creation. Existing methods…