collaborators

10 papers

cs.CV2026

Recompute or Reuse? Diagnosing and Mitigating Textual Shortcuts in VLM Self-Reflection

Wenxiao Fan, Jingling Fu, Fang Li +9

Vision-language models (VLMs) are expected to revise their reasoning when visual evidence changes. Failures to do so are often attributed to insufficient visual attention or contex…

cs.CV2026

Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On

Yong Liu, Xiaolong Fu, Zihang Xu +10

We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for t…

cs.CV2026

GMO-EDIT: Grounded Multi-Operation Editing for E-Commerce Images

Zipeng Guo, Xiaoan Liu, Lichen Ma +9

Real-world e-commerce image editing often requires multiple, localized, and auditable operations rather than global restyling. This compositional nature poses a dual challenge: mod…

cs.CV2026

PixelU: A U-Shaped Transformer for Efficient End-to-End Pixel Diffusion

Zipeng Guo, Lichen Ma, Yu He +4

End-to-end pixel-space diffusion models bypass the lossy compression of Latent Diffusion Models (LDMs) but struggle to jointly model low-frequency semantics and high-frequency sign…

cs.CV2026

HyperDiT: Hyper-Connected Transformers for High-Fidelity Pixel-Space Diffusion

Yu He, Lichen Ma, Zipeng Guo +5

Pixel-space diffusion models bypass the reconstruction bottleneck of Variational Autoencoders (VAEs) but face a fundamental "granularity dilemma": capturing global semantics favors…

cs.CV2026

FrequencyBooster: Full-Frequency Modeling for High-Fidelity Pixel Diffusion

Lichen Ma, Zipeng Guo, Yu He +5

To circumvent the inherent fidelity bottlenecks and optimization misalignment of VAE-based latent diffusion, pixel-space diffusion models have emerged as a compelling end-to-end pa…