4 papers
iDiff: Interpretable Difference-aware Framework for Pairwise Image Quality Assessment
Xinli Yue, JianHui Sun, Tao Shao +3
Pairwise image quality assessment (IQA) in professional photography requires a model not only to identify the preferred image between two candidates, but also to provide convincing…
TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens
Jianpeng Cheng, Xian Wu, Jiangfan Zhang +10
Recent research has demonstrated that Universal Multimodal Embedding (UME) benefits significantly from Chain-of-Thought (CoT) reasoning. In this paradigm, a generative model produc…
iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA
Zhaoran Zhao, Xinli Yue, Jianhui Sun +5
Image Quality Assessment (IQA) has progressed from scalar quality prediction to more interpretable, human-aligned evaluation paradigms. In this work, we address the emerging challe…
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
Xinli Yue, JianHui Sun, Junda Lu +6
With the rapid advancement of text-to-image (T2I) generation models, assessing the semantic alignment between generated images and text descriptions has become a significant resear…