2 papers
cs.CV2026
Model the Edit, Not the Image: Visual Autoregressive Editing from a Source-Centric Perspective
Hongyi Fang, Chuwen Xie, Benjia Zhou +6
Next-scale visual autoregressive models (VARs) have emerged as a powerful generative paradigm, producing high-quality images through efficient coarse-to-fine prediction. However, t…
cs.CV2026
SPAE: Spectrally Guided Autoencoder for Pretrained Visual Latents
Yibin Huang, Jixiang Hong, Zongzhao Li +8
Latents from vision foundation models (VFMs) are semantically rich and well suited for visual understanding. Recent representation autoencoder methods such as RAE have shown that t…