Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Progressive Compositionality in Text-to-Image Generative Models
Evans Xu Han, Linghao Jin, Xiaofeng Liu +1
Despite the impressive text-to-image (T2I) synthesis capabilities of diffusion models, they often struggle to understand compositional relationships between objects and attributes,…
cs.CV2025
Light-weight Fine-tuning Method for Defending Adversarial Noise in Pre-trained Medical Vision-Language Models
Xu Han, Linghao Jin, Xuezhe Ma +1
Fine-tuning pre-trained Vision-Language Models (VLMs) has shown remarkable capabilities in medical image and textual depiction synergy. Nevertheless, many pre-training datasets are…
cs.CV2025
Fair Text to Medical Image Diffusion Model with Subgroup Distribution Aligned Tuning
Xu Han, Fangfang Fan, Jingzhao Rong +4
The text to medical image (T2MedI) with latent diffusion model has great potential to alleviate the scarcity of medical imaging data and explore the underlying appearance distribut…