3 papers
cs.CV2025
From Sequential to Spatial: Reordering Autoregression for Efficient Visual Generation
Siyang Wang, Hanting Li, Wei Li +3
Inspired by the remarkable success of autoregressive models in language modeling, this paradigm has been widely adopted in visual generation. However, the sequential token-by-token…
cs.CV2025
Beyond Textual CoT: Interleaved Text-Image Chains with Deep Confidence Reasoning for Image Editing
Zhentao Zou, Zhengrong Yue, Kunpeng Du +9
Image editing with natural language has gained significant popularity, yet existing methods struggle with intricate object intersections and fine-grained spatial relationships due…
eess.IV2025
RDDM: Practicing RAW Domain Diffusion Model for Real-world Image Restoration
Yan Chen, Yi Wen, Wei Li +4
We present the RAW domain diffusion model (RDDM), an end-to-end diffusion model that restores photo-realistic images directly from the sensor RAW data. While recent sRGB-domain dif…