3 papers
cs.CV2026
PixelU: A U-Shaped Transformer for Efficient End-to-End Pixel Diffusion
Zipeng Guo, Lichen Ma, Yu He +4
End-to-end pixel-space diffusion models bypass the lossy compression of Latent Diffusion Models (LDMs) but struggle to jointly model low-frequency semantics and high-frequency sign…
cs.LG2026
Meta-TTRL: A Metacognitive Framework for Self-Improving Test-Time Reinforcement Learning in Unified Multimodal Models
Lit Sin Tan, Junzhe Chen, Xiaolong Fu +5
Existing test-time scaling (TTS) methods for unified multimodal models (UMMs) in text-to-image (T2I) generation primarily rely on search or sampling strategies that produce only in…
cs.CV2025
RePainter: Empowering E-commerce Object Removal via Spatial-matting Reinforcement Learning
Zipeng Guo, Lichen Ma, Xiaolong Fu +13
In web data, product images are central to boosting user engagement and advertising efficacy on e-commerce platforms, yet the intrusive elements such as watermarks and promotional…