collaborators

7 papers

cs.CV2026

iFAN: Inference-Aware Learning for Plain Mask Transformers

Fang Li, Yu He, Haoyang Tong +7

Query-based mask transformers assemble segmentation outputs through pixel-wise competition among query predictions of the final layer, yet this inference process is not explicitly…

cs.CV2026

Recompute or Reuse? Diagnosing and Mitigating Textual Shortcuts in VLM Self-Reflection

Wenxiao Fan, Jingling Fu, Fang Li +9

Vision-language models (VLMs) are expected to revise their reasoning when visual evidence changes. Failures to do so are often attributed to insufficient visual attention or contex…

cs.CV2026

GMO-EDIT: Grounded Multi-Operation Editing for E-Commerce Images

Zipeng Guo, Xiaoan Liu, Lichen Ma +9

Real-world e-commerce image editing often requires multiple, localized, and auditable operations rather than global restyling. This compositional nature poses a dual challenge: mod…

cs.CV2026

LiWi: Layering in the Wild

Yu He, Fang Li, Haoyang Tong +7

Recent advances in generative models have empowered impressive layered image generation, yet their success is largely confined to graphic design domains. The layering of in-the-wil…

cs.CV2026

FrequencyBooster: Full-Frequency Modeling for High-Fidelity Pixel Diffusion

Lichen Ma, Zipeng Guo, Yu He +5

To circumvent the inherent fidelity bottlenecks and optimization misalignment of VAE-based latent diffusion, pixel-space diffusion models have emerged as a compelling end-to-end pa…

cs.CV2025

NTIRE 2025 XGC Quality Assessment Challenge: Methods and Results

Xiaohong Liu, Xiongkuo Min, Qiang Hu +92

This paper reports on the NTIRE 2025 XGC Quality Assessment Challenge, which will be held in conjunction with the New Trends in Image Restoration and Enhancement Workshop (NTIRE) a…