6 papers
MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation
Shichao Kan, Xuyang Zhang, Haojie Zhang +7
Evaluating image captions without references remains challenging because global embedding similarity often misses fine-grained mismatches such as hallucinated objects, missing attr…
V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval
Dongyang Chen, Chaoyang Wang, Dezhao Su +6
Multimodal Large Language Models (MLLMs) have recently been applied to universal multimodal retrieval, where Chain-of-Thought (CoT) reasoning improves candidate reranking. However,…
Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental Learning
Haojie Zhang, Yixiong Liang, Hulin Kuang +5
Multimodal Biomedical Image Incremental Learning (MBIIL) is essential for handling diverse tasks and modalities in the biomedical domain, as training separate models for each modal…
TMCIR: Token Merge Benefits Composed Image Retrieval
Chaoyang Wang, Zeyu Zhang, Long Teng +2
Composed Image Retrieval (CIR) retrieves target images using a multi-modal query that combines a reference image with text describing desired modifications. The primary challenge i…
HRDecoder: High-Resolution Decoder Network for Fundus Image Lesion Segmentation
Ziyuan Ding, Yixiong Liang, Shichao Kan +1
High resolution is crucial for precise segmentation in fundus images, yet handling high-resolution inputs incurs considerable GPU memory costs, with diminishing performance gains a…
How Does the Smoothness Approximation Method Facilitate Generalization for Federated Adversarial Learning?
Wenjun Ding, Ying An, Lixing Chen +3
Federated Adversarial Learning (FAL) is a robust framework for resisting adversarial attacks on federated learning. Although some FAL studies have developed efficient algorithms, t…