19 papers
Beyond Token-Level Cross-Entropy: Fréchet Distributional Post-Training for Autoregressive Image Generation
Jinhua Zhang, Yisong Lin, Wei Long +1
Autoregressive image generators are commonly pretrained with token-level cross-entropy under teacher forcing, yet evaluated by the distributional quality of decoded images. This cr…
Perceiving Better Moments: Cover Frame Reselection and Enhancement for Live Photos with the Live2K Dataset
Junyu Lou, Kai Chen, Weiyi You +3
Modern smartphones capture Live Photos, short video bursts surrounding a still image, offering a dynamic and engaging photographic experience. However, the cover photo and video co…
From Local Windows to Adaptive Candidates via Individualized Exploratory: Rethinking Attention for Image Super-Resolution
Chunyu Meng, Wei Long, Shuhang Gu
Single Image Super-Resolution (SISR) is a fundamental computer vision task that aims to reconstruct a high-resolution (HR) image from a low-resolution (LR) input. Transformer-based…
Neural Stereo Video Compression with Hybrid Disparity Compensation
Shiyin Jiang, Zhenghao Chen, Minghao Han +1
Disparity compensation represents the primary strategy in stereo video compression (SVC) for exploiting cross-view redundancy. These mechanisms can be broadly categorized into two…
Guiding a Diffusion Transformer with the Internal Dynamics of Itself
Xingyu Zhou, Qifan Li, Xiaobin Hu +2
The diffusion model presents a powerful ability to capture the entire (conditional) data distribution. However, due to the lack of sufficient training and data to learn to cover lo…
Taming Sampling Perturbations with Variance Expansion Loss for Latent Diffusion Models
Qifan Li, Xingyu Zhou, Jinhua Zhang +2
Latent diffusion models have emerged as the dominant framework for high-fidelity and efficient image generation, owing to their ability to learn diffusion processes in compact late…