clinical classification 1colonoscopy 1image-text retrieval 1report grounding 1vision-language models 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.AI2026
A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
Jia Yu, Yan Zhu, Yili He +12
The paper presents EndoCLIP, a vision‑language foundation model for colonoscopy that learns from lesion‑level image‑text pairs extracted from routine colonoscopy reports, achieving…
cs.CV2025
Endo-CLIP: Progressive Self-Supervised Pre-training on Raw Colonoscopy Records
Yili He, Yan Zhu, Peiyao Fu +7
Pre-training on image-text colonoscopy records offers substantial potential for improving endoscopic image analysis, but faces challenges including non-informative background image…