4 papers · 1 filter
TP-Blend: Textual-Prompt Attention Pairing for Precise Object-Style Blending in Diffusion Models
Xin Jin, Yichuan Zhong, Yapeng Tian
Current text-conditioned diffusion editors handle single object replacement well but struggle when a new object and a new style must be introduced simultaneously. We present Twin-P…
Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield
Dongyang Liu, Peng Gao, David Liu +8
Diffusion model distillation has emerged as a powerful technique for creating efficient few-step and single-step generators. Among these, Distribution Matching Distillation (DMD) a…
AesTest: Measuring Aesthetic Intelligence from Perception to Production
Guolong Wang, Heng Huang, Zhiqiang Zhang +3
Perceiving and producing aesthetic judgments is a fundamental yet underexplored capability for multimodal large language models (MLLMs). However, existing benchmarks for image aest…
Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
Yunhao Gou, Kai Chen, Zhili Liu +5
Recent breakthroughs in reasoning language models have significantly advanced text-based reasoning. On the other hand, Multi-modal Large Language Models (MLLMs) still lag behind, h…