5 papers
Drift-AR: Single-Step Visual Autoregressive Generation via Anti-Symmetric Drifting
Zhen Zou, Xiaoxiao Ma, Mingde Yao +3
Autoregressive (AR)-Diffusion hybrid paradigms combine AR's structured semantic modeling with diffusion's high-fidelity synthesis, yet suffer from a dual speed bottleneck: the sequ…
OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation
Guohui Zhang, XiaoXiao Ma, Jie Huang +9
Recent advances in joint audio-video generation have been remarkable, yet real-world applications demand strong per-modality fidelity, cross-modal alignment, and fine-grained synch…
RAE-AR: Taming Autoregressive Models with Representation Autoencoders
Hu Yu, Hang Xu, Jie Huang +4
The latent space of generative modeling is long dominated by the VAE encoder. The latents from the pretrained representation encoders (e.g., DINO, SigLIP, MAE) are previously consi…
Highly Efficient Test-Time Scaling for T2I Diffusion Models with Text Embedding Perturbation
Hang Xu, Linjiang Huang, Feng Zhao
Test-time scaling (TTS) aims to achieve better results by increasing random sampling and evaluating samples based on rules and metrics. However, in text-to-image(T2I) diffusion mod…
FR-TTS: Test-Time Scaling for NTP-based Image Generation with Effective Filling-based Reward Signal
Hang Xu, Linjiang Huang, Feng Zhao
Test-time scaling (TTS) has become a prevalent technique in image generation, significantly boosting output quality by expanding the number of parallel samples and filtering them u…