From the 2 of 11 linked papers with an AI index.
8 papers · 1 filter
Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion
Henglin Liu, Fangyuan Kong, Jing Wang +7
The paper introduces concentrated Implicit Preference Optimization (cIPO), a post‑training method for text‑to‑video diffusion models that derives preference signals from reconstruc…
MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control
Kaiqi Liu, Yunyao Mao, Ziqi Cai +8
MAVIN is a framework for generating multi-shot audio‑visual content with fine‑grained narrative control, using boundary‑aware attention to align temporal segments and ID‑aware prop…
HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
Qi Cai, Jingwen Chen, Chengmin Gao +22
The evolution of visual generative models has long been constrained by fragmented architectures relying on disjoint text encoders and external VAEs. In this report, we present HiDr…
Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping
Haoyuan Sun, Jing Wang, Yuxin Song +9
Recently, post-training methods based on reinforcement learning, with a particular focus on Group Relative Policy Optimization (GRPO), have emerged as the robust paradigm for furth…
CADC: Content Adaptive Diffusion-Based Generative Image Compression
Xihua Sheng, Lingyu Zhu, Tianyu Zhang +3
Diffusion-based generative image compression has demonstrated remarkable potential for achieving realistic reconstruction at ultra-low bitrates. The key to unlocking this potential…
DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO
Henglin Liu, Huijuan Huang, Jing Wang +3
Reinforcement learning (RL), particularly GRPO, improves image generation quality significantly by comparing the relative performance of images generated within the same group. How…