3 papers
cs.CV2025
X2Edit: Revisiting Arbitrary-Instruction Image Editing through Self-Constructed Data and Task-Aware Representation Learning
Jian Ma, Xujie Zhu, Zihao Pan +4
Existing open-source datasets for arbitrary-instruction image editing remain suboptimal, while a plug-and-play editing module compatible with community-prevalent generative models…
cs.CV2025
X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation
Jian Ma, Qirong Peng, Xu Guo +3
Text-to-image (T2I) models are well known for their ability to produce highly realistic images, while multimodal large language models (MLLMs) are renowned for their proficiency in…
cs.LG2025
SchoenbAt: Rethinking Attention with Polynomial basis
Yuhan Guo, Lizhong Ding, Yuwan Yang +1
Kernelized attention extends the attention mechanism by modeling sequence correlations through kernel functions, making significant progresses in optimizing attention. Under the gu…