6 papers
iDPA: Instance Decoupled Prompt Attention for Incremental Medical Object Detection
Huahui Yi, Wei Xu, Ziyuan Qin +4
Existing prompt-based approaches have demonstrated impressive performance in continual learning, leveraging pre-trained large-scale models for classification tasks; however, the ti…
Confounder-Aware Medical Data Selection for Fine-Tuning Pretrained Vision Models
Anyang Ji, Qingbo Kang, Wei Xu +3
The emergence of large-scale pre-trained vision foundation models has greatly advanced the medical imaging field through the pre-training and fine-tuning paradigm. However, selecti…
Guiding Medical Vision-Language Models with Explicit Visual Prompts: Framework Design and Comprehensive Exploration of Prompt Variations
Kangyu Zhu, Ziyuan Qin, Huahui Yi +4
While mainstream vision-language models (VLMs) have advanced rapidly in understanding image level information, they still lack the ability to focus on specific areas designated by…
One-to-Normal: Anomaly Personalization for Few-shot Anomaly Detection
Yiyue Li, Shaoting Zhang, Kang Li +1
Traditional Anomaly Detection (AD) methods have predominantly relied on unsupervised learning from extensive normal data. Recent AD methods have evolved with the advent of large pr…
Evaluating Hallucination in Text-to-Image Diffusion Models with Scene-Graph based Question-Answering Agent
Ziyuan Qin, Dongjie Cheng, Haoyu Wang +5
Contemporary Text-to-Image (T2I) models frequently depend on qualitative human evaluations to assess the consistency between synthesized images and the text prompts. There is a dem…
TV-SAM: Increasing Zero-Shot Segmentation Performance on Multimodal Medical Images Using GPT-4 Generated Descriptive Prompts Without Human Annotation
Zekun Jiang, Dongjie Cheng, Ziyuan Qin +10
This study presents a novel multimodal medical image zero-shot segmentation algorithm named the text-visual-prompt segment anything model (TV-SAM) without any manual annotations. T…