2 papers
cs.CV2025
TextSR: Diffusion Super-Resolution with Multilingual OCR Guidance
Keren Ye, Ignacio Garcia Dorado, Michalis Raptis +4
While recent advancements in Image Super-Resolution (SR) using diffusion models have shown promise in improving overall image quality, their application to scene text images has re…
cs.CV2025
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
Lijie Fan, Luming Tang, Siyang Qin +11
We present UniFluid, a unified autoregressive framework for joint visual generation and understanding leveraging continuous visual tokens. Our unified autoregressive architecture p…