collaborators

5 papers

cs.CV2025

DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation

Jiawei Liu, Junqiao Li, Jiangfan Deng +11

The "one-shot" technique represents a distinct and sophisticated aesthetic in filmmaking. However, its practical realization is often hindered by prohibitive costs and complex real…

cs.CV2025

Better Supervised Fine-tuning for VQA: Integer-Only Loss

Baihong Qian, Haotian Fan, Wenjie Liao +3

With the rapid advancement of vision language models(VLM), their ability to assess visual content based on specific criteria and dimensions has become increasingly critical for app…

cs.CV2025

NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment

Shuhao Han, Haotian Fan, Fangyuan Kong +112

This paper reports on the NTIRE 2025 challenge on Text to Image (T2I) generation model quality assessment, which will be held in conjunction with the New Trends in Image Restoratio…

cs.CV2025

Grounded Knowledge-Enhanced Medical Vision-Language Pre-training for Chest X-Ray

Qiao Deng, Zhongzhen Huang, Yunqi Wang +6

Medical foundation models have the potential to revolutionize healthcare by providing robust and generalized representations of medical data. Medical vision-language pre-training h…

cs.CV2024

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Shuhao Han, Haotian Fan, Jiachen Fu +8

Recently, Text-to-Image (T2I) generation models have achieved significant advancements. Correspondingly, many automated metrics have emerged to evaluate the image-text alignment ca…