5 papers
FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling
Guixian Xu, Yide Liang, Zeli Su +5
Vision-language models have progressed rapidly, but Tibetan remains a severely underserved low-resource language due to the lack of reproducible training and evaluation infrastruct…
InEdit-Bench: Benchmarking Intermediate Logical Pathways for Intelligent Image Editing Models
Zhiqiang Sheng, Xumeng Han, Zhiwei Zhang +6
Multimodal generative models have made significant strides in image editing, demonstrating impressive performance on a variety of static tasks. However, their proficiency typically…
Progressive Compositionality in Text-to-Image Generative Models
Evans Xu Han, Linghao Jin, Xiaofeng Liu +1
Despite the impressive text-to-image (T2I) synthesis capabilities of diffusion models, they often struggle to understand compositional relationships between objects and attributes,…
Light-weight Fine-tuning Method for Defending Adversarial Noise in Pre-trained Medical Vision-Language Models
Xu Han, Linghao Jin, Xuezhe Ma +1
Fine-tuning pre-trained Vision-Language Models (VLMs) has shown remarkable capabilities in medical image and textual depiction synergy. Nevertheless, many pre-training datasets are…
Fair Text to Medical Image Diffusion Model with Subgroup Distribution Aligned Tuning
Xu Han, Fangfang Fan, Jingzhao Rong +4
The text to medical image (T2MedI) with latent diffusion model has great potential to alleviate the scarcity of medical imaging data and explore the underlying appearance distribut…