3 papers
cs.CV2026
Training-Free Semantic Correction for Autoregressive Visual Models
Junhao Chen, Chanyu Zhu, Zheqi Lv +2
Autoregressive visual models (AVMs) based on next-scale prediction have emerged as a prominent paradigm for image and video synthesis. However, decomposing the generation process i…
cs.CV2025
MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation
Junhao Chen, Yulia Tsvetkov, Xiaochuang Han
Recent progress in multimodal generation has increasingly combined autoregressive (AR) and diffusion-based approaches, leveraging their complementary strengths: AR models capture l…
cs.CV2025
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
Zheqi Lv, Junhao Chen, Qi Tian +3
Diffusion models have become the mainstream architecture for text-to-image generation, achieving remarkable progress in visual quality and prompt controllability. However, current…