9 papers
GMO-EDIT: Grounded Multi-Operation Editing for E-Commerce Images
Zipeng Guo, Xiaoan Liu, Lichen Ma +9
Real-world e-commerce image editing often requires multiple, localized, and auditable operations rather than global restyling. This compositional nature poses a dual challenge: mod…
PixelU: A U-Shaped Transformer for Efficient End-to-End Pixel Diffusion
Zipeng Guo, Lichen Ma, Yu He +4
End-to-end pixel-space diffusion models bypass the lossy compression of Latent Diffusion Models (LDMs) but struggle to jointly model low-frequency semantics and high-frequency sign…
HyperDiT: Hyper-Connected Transformers for High-Fidelity Pixel-Space Diffusion
Yu He, Lichen Ma, Zipeng Guo +5
Pixel-space diffusion models bypass the reconstruction bottleneck of Variational Autoencoders (VAEs) but face a fundamental "granularity dilemma": capturing global semantics favors…
FrequencyBooster: Full-Frequency Modeling for High-Fidelity Pixel Diffusion
Lichen Ma, Zipeng Guo, Yu He +5
To circumvent the inherent fidelity bottlenecks and optimization misalignment of VAE-based latent diffusion, pixel-space diffusion models have emerged as a compelling end-to-end pa…
Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling
Xiaolong Fu, Lichen Ma, Zipeng Guo +9
The integration of Reinforcement Learning (RL) into flow matching models for text-to-image (T2I) generation has driven substantial advances in generation quality. However, these ga…
UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
Lichen Ma, Xiaolong Fu, Gaojing Zhou +6
With the rapid advancement of image generation, visual text editing using natural language instructions has received increasing attention. The main challenge of this task is to ful…