AttriPrompter: Auto-Prompting with Attribute Semantics for Zero-shot Nuclei Detection via Visual-Language Pre-trained Models
arXiv:2410.16820 · doi:10.1109/TMI.2024.3473745
Abstract
Large-scale visual-language pre-trained models (VLPMs) have demonstrated exceptional performance in downstream object detection through text prompts for natural scenes. However, their application to zero-shot nuclei detection on histopathology images remains relatively unexplored, mainly due to the significant gap between the characteristics of medical images and the web-originated text-image pairs used for pre-training. This paper aims to investigate the potential of the object-level VLPM, Grounded Language-Image Pre-training (GLIP), for zero-shot nuclei detection. Specifically, we propose an innovative auto-prompting pipeline, named AttriPrompter, comprising attribute generation, attribute augmentation, and relevance sorting, to avoid subjective manual prompt design. AttriPrompter utilizes VLPMs' text-to-image alignment to create semantically rich text prompts, which are then fed into GLIP for initial zero-shot nuclei detection. Additionally, we propose a self-trained knowledge distillation framework, where GLIP serves as the teacher with its initial predictions used as pseudo labels, to address the challenges posed by high nuclei density, including missed detections, false positives, and overlapping instances. Our method exhibits remarkable performance in label-free nuclei detection, outperforming all existing unsupervised methods and demonstrating excellent generality. Notably, this work highlights the astonishing potential of VLPMs pre-trained on natural image-text pairs for downstream tasks in the medical field as well. Code will be released at https://github.com/wuyongjianCODE/AttriPrompter.
This article has been accepted for publication in a future issue of IEEE Transactions on Medical Imaging (TMI), but has not been fully edited. Content may change prior to final publication. Citation information: DOI: https://doi.org/10.1109/TMI.2024.3473745 . Code: https://github.com/wuyongjianCODE/AttriPrompter
References in corpus (21)
- Distilling the Knowledge in a Neural Network
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Learning to Prompt for Vision-Language Models
- YOLOX: Exceeding YOLO Series in 2021
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks
- CLIP-Driven Universal Model for Organ Segmentation and Tumor Detection
- VLP: A Survey on Vision-Language Pre-training
- Automatic Chain of Thought Prompting in Large Language Models
- Weakly Supervised Deep Nuclei Segmentation Using Partial Points Annotation in Histopathology Images
- Medical-VLBERT: Medical Visual Language BERT for COVID-19 CT Report Generation With Alternate Learning
- A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models
- Self-Supervised Nuclei Segmentation in Histopathological Images Using Attention
- Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?
- Vision-Language Models for Vision Tasks: A Survey
- Learning Object-Language Alignments for Open-Vocabulary Object Detection
- Medical Image Understanding with Pretrained Vision Language Models: A Comprehensive Study
- Cyclic Learning: Bridging Image-level Labels and Nuclei Instance Segmentation
- Signet Ring Cell Detection With a Semi-supervised Learning Framework
- When are Lemons Purple? The Concept Association Bias of Vision-Language Models
- Unsupervised Dense Nuclei Detection and Segmentation with Prior Self-activation Map For Histology Images