5 papers
VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation
Xinyao Liao, Qiyuan He, Yicong Li +4
Autoregressive image and video generators are trained with teacher-forced histories but must sample from their own generated prefixes at inference time, making them vulnerable to e…
VA-: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
Xinyao Liao, Qiyuan He, Kai Xu +4
Autoregressive (AR) visual generation relies on tokenizers to map images to and from discrete sequences. However, tokenizers are trained to reconstruct clean images from ground-tru…
REAR: Rethinking Visual Autoregressive Models via Generator-Tokenizer Consistency Regularization
Qiyuan He, Yicong Li, Haotian Ye +6
Visual autoregressive (AR) generation offers a promising path toward unifying vision and language models, yet its performance remains suboptimal against diffusion models. Prior wor…
CoCA: Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning
Xinyao Liao, Wei Wei, Xiaoye Qu +3
Recent advances in text-to-image (T2I) diffusion model fine-tuning leverage reinforcement learning (RL) to align generated images with learnable reward functions. The existing appr…
UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph Generation
Xinyao Liao, Wei Wei, Dangyang Chen +1
Scene Graph Generation(SGG) is a scene understanding task that aims at identifying object entities and reasoning their relationships within a given image. In contrast to prevailing…