collaborators

5 papers

cs.CV2026

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling

Guixian Xu, Yide Liang, Zeli Su +5

Vision-language models have progressed rapidly, but Tibetan remains a severely underserved low-resource language due to the lack of reproducible training and evaluation infrastruct…

cs.CV2026

InEdit-Bench: Benchmarking Intermediate Logical Pathways for Intelligent Image Editing Models

Zhiqiang Sheng, Xumeng Han, Zhiwei Zhang +6

Multimodal generative models have made significant strides in image editing, demonstrating impressive performance on a variety of static tasks. However, their proficiency typically…

cs.CV2025

Progressive Compositionality in Text-to-Image Generative Models

Evans Xu Han, Linghao Jin, Xiaofeng Liu +1

Despite the impressive text-to-image (T2I) synthesis capabilities of diffusion models, they often struggle to understand compositional relationships between objects and attributes,…

cs.CV2025

Light-weight Fine-tuning Method for Defending Adversarial Noise in Pre-trained Medical Vision-Language Models

Xu Han, Linghao Jin, Xuezhe Ma +1

Fine-tuning pre-trained Vision-Language Models (VLMs) has shown remarkable capabilities in medical image and textual depiction synergy. Nevertheless, many pre-training datasets are…

cs.CV2025

Fair Text to Medical Image Diffusion Model with Subgroup Distribution Aligned Tuning

Xu Han, Fangfang Fan, Jingzhao Rong +4

The text to medical image (T2MedI) with latent diffusion model has great potential to alleviate the scarcity of medical imaging data and explore the underlying appearance distribut…