6 papers
MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing
Zitong Xu, Huiyu Duan, Xinyun Zhang +7
Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-Pro and GPT-Image-2 demonstr…
Post-training makes large language models less human-like
Marcel Binz, Elif Akata, Abdullah Almaatouq +76
Large language models (LLMs) are increasingly used as surrogates for human participants, but it remains unclear which models best capture human behavior and why. To address this, w…
ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning
Honghua Chen, Zitong Xu, Huiyu Duan +3
Recent text-guided image editing (TIE) models have achieved remarkable progress, however, many edited results still suffer from artifacts, unintended modifications, and suboptimal…
Learning Visual Feature-Based World Models via Residual Latent Action
Xinyu Zhang, Zhengtong Xu, Yutian Tao +3
World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature-based world models, on the other…
Saliency-Aware Regularized Quantization Calibration for Large Language Models
Yanlong Zhao, Xiaoyuan Cheng, Huihang Liu +6
Post-training quantization (PTQ) is an effective approach for deploying large language models (LLMs) under memory and latency constraints. Most existing PTQ methods determine quant…
Plug-and-play Class-aware Knowledge Injection for Prompt Learning with Visual-Language Model
Junhui Yin, Nan Pu, Xinyu Zhang +4
Prompt learning has become an effective and widely used technique in enhancing vision-language models (VLMs) such as CLIP for various downstream tasks, particularly in zero-shot cl…