collaborators

14 papers

cs.CV2026

Energy-Guided Flow Matching

Haoyang Tong, Yu He, Fang Li +6

Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine-grained details in a high-dimensional space. Standard flo…

cs.CV2026

iFAN: Inference-Aware Learning for Plain Mask Transformers

Fang Li, Yu He, Haoyang Tong +7

Query-based mask transformers assemble segmentation outputs through pixel-wise competition among query predictions of the final layer, yet this inference process is not explicitly…

cs.CV2026

Recompute or Reuse? Diagnosing and Mitigating Textual Shortcuts in VLM Self-Reflection

Wenxiao Fan, Jingling Fu, Fang Li +9

Vision-language models (VLMs) are expected to revise their reasoning when visual evidence changes. Failures to do so are often attributed to insufficient visual attention or contex…

cs.CV2026

GMO-EDIT: Grounded Multi-Operation Editing for E-Commerce Images

Zipeng Guo, Xiaoan Liu, Lichen Ma +9

Real-world e-commerce image editing often requires multiple, localized, and auditable operations rather than global restyling. This compositional nature poses a dual challenge: mod…

cs.CV2026

PixelU: A U-Shaped Transformer for Efficient End-to-End Pixel Diffusion

Zipeng Guo, Lichen Ma, Yu He +4

End-to-end pixel-space diffusion models bypass the lossy compression of Latent Diffusion Models (LDMs) but struggle to jointly model low-frequency semantics and high-frequency sign…

cs.LG2026

Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning

Hu Xu, Zhaolong Xing, Congcong Liu +5

Calibration data are often treated as a minor implementation detail in post-training LLM pruning because averaged evaluations suggest only modest effects. We show that this conclus…