4 papers · 1 filter
HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs
Zijian Zhang, Xuecheng Wu, Danlei Huang +3
Driven by the rapid progress in vision-language models (VLMs), the responsible behavior of large-scale multimodal models has become a prominent research area, particularly focusing…
NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment
Shuhao Han, Haotian Fan, Fangyuan Kong +112
This paper reports on the NTIRE 2025 challenge on Text to Image (T2I) generation model quality assessment, which will be held in conjunction with the New Trends in Image Restoratio…
TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs
Zijian Zhang, Xuhui Zheng, Xuecheng Wu +2
While text-to-image (T2I) generation models have achieved remarkable progress in recent years, existing evaluation methodologies for vision-language alignment still struggle with t…
Fuse after Align: Improving Face-Voice Association Learning via Multimodal Encoder
Chong Peng, Liqiang He, Dan Su
Today, there have been many achievements in learning the association between voice and face. However, most previous work models rely on cosine similarity or L2 distance to evaluate…