medical image analysis

DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis

arXiv:2607.27763

summary

The paper details DS@GT's approaches for the ImageCLEFmedical 2026 challenge, using a late‑fusion ensemble of ConvNeXt‑V2, BiomedCLIP, and DenseNet‑169 with an Honest Threshold Tuning method for radiology concept detection, and various foundation‑model based pipelines for generating medical image captions.

Abstract

We describe the DS@GT submissions to the ImageCLEFmedical Caption 2026 challenge, which continues a long-running benchmark on the ROCOv2 dataset with two tracks: Concept Detection (Task 1), assigning UMLS Concept Unique Identifiers (CUIs) to radiology images, and Caption Prediction (Task 2), generating natural-language captions. For Task 1, our primary submission was a three-way late-fusion ensemble of ConvNeXt-V2, BiomedCLIP ViT-B/16, and DenseNet-169 with a regularized ''Honest Threshold Tuning'' procedure designed to avoid validation overfitting on rare concepts; this submission ranked first on the official submission with a primary of and a secondary of . In parallel, we submitted a training-free KNN retrieval pipeline over frozen BiomedCLIP embeddings, which reached a primary of and a secondary of -essentially matching the fine-tuned ensemble on the primary track at a fraction of the cost. For Task 2, our submissions included a fine-tuned Gemma-3 27B model (overall , ranking third in the official submission), a fully fine-tuned BLIP pipeline with custom Vizwins merging (), and a zero-shot MedGemma-4B run with a PubMed-style prompt (), spanning a wide range of model scales and training costs. Code: https://github.com/dsgt-arc/imageclef-caption-2026.

21 pages, 9 figures

Topics & keywords

#concept detection#medical image captioning#ensemble learning#foundation models#radiologyConvNeXt-V2BiomedCLIPDenseNet-169Honest Threshold TuningKNN retrievalGemma-3BLIPUMLS CUIROCOv2
DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis · wovepaper