6 papers · 1 filter
Pixel-Space Diffusion via Observation Operators
Shaojie Guo, Lichen Ma, Haoyang Tong +8
Pixel-space diffusion models directly model image distributions but remain difficult to optimize. Recent methods alleviate this challenge through target reparameterization, while s…
PosterText: Towards Unified Visual Text Generation and Editing for E-commerce Poster
Xiaoan Liu, Lichen Ma, Zipeng Guo +12
Automated e-commerce poster design requires both high-quality poster generation and flexible editing of existing designs. However, most existing methods either target end-to-end po…
TransAnyText: Translating Arbitrary Text in E-commerce Images via Structured Visual Generation
Xiaoan Liu, Lichen Ma, Zipeng Guo +13
Cross-border e-commerce image translation is essential for global retail, where product images, banners, and detail pages need to be produced in different languages. Existing metho…
Energy-Guided Flow Matching
Haoyang Tong, Yu He, Fang Li +6
Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine-grained details in a high-dimensional space. Standard flo…
iFAN: Inference-Aware Learning for Plain Mask Transformers
Fang Li, Yu He, Haoyang Tong +7
Query-based mask transformers assemble segmentation outputs through pixel-wise competition among query predictions of the final layer, yet this inference process is not explicitly…
Recompute or Reuse? Diagnosing and Mitigating Textual Shortcuts in VLM Self-Reflection
Wenxiao Fan, Jingling Fu, Fang Li +9
Vision-language models (VLMs) are expected to revise their reasoning when visual evidence changes. Failures to do so are often attributed to insufficient visual attention or contex…