6 papers
A Testable Certificate for Constant Collapse in Teacher-Guided VAEs
Zegu Zhang, Jianhua Peng, Jian Zhang
Posterior collapse in variational autoencoders is often diagnosed by its symptoms: a small KL term, a strong decoder, or weak use of the latent code. These signals are useful, but…
StableI2I: Spotting Unintended Changes in Image-to-Image Transition
Jiayang Li, Shuo Cao, Xiaohui Li +6
In most real-world image-to-image (I2I) scenarios, existing evaluations primarily focus on instruction following and the perceptual quality or aesthetics of the generated images. H…
EditHF-1M: A Million-Scale Rich Human Preference Feedback for Image Editing
Zitong Xu, Huiyu Duan, Zhongpeng Ji +9
Recent text-guided image editing (TIE) models have achieved remarkable progress, while many edited images still suffer from issues such as artifacts, unexpected editings, unaesthet…
OARS: Process-Aware Online Alignment for Generative Real-World Image Super-Resolution
Shijie Zhao, Xuanyu Zhang, Bin Chen +8
Aligning generative real-world image super-resolution models with human visual preference is challenging due to the perception--fidelity trade-off and diverse, unknown degradations…
Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World
Yuzhi Huang, Kairun Wen, Rongxin Gao +14
Humans inhabit a physical 4D world where geometric structure and semantic content evolve over time, constituting a dynamic 4D reality (spatial with temporal dimension). While curre…
Text-to-Image GAN with Pretrained Representations
Xiaozhou You, Jian Zhang
Generating desired images conditioned on given text descriptions has received lots of attention. Recently, diffusion models and autoregressive models have demonstrated their outsta…