2 papers
cs.CV2026
Gastric-X: A Multimodal Multi-Phase Benchmark Dataset for Advancing Vision-Language Models in Gastric Cancer Analysis
Sheng Lu, Hao Chen, Rui Yin +3
Recent vision-language models (VLMs) have shown strong generalization and multimodal reasoning abilities in natural domains. However, their application to medical diagnosis remains…
cs.CV2025
AVPDN: Learning Motion-Robust and Scale-Adaptive Representations for Video-Based Polyp Detection
Zilin Chen, Shengnan Lu
Accurate detection of polyps is of critical importance for the early and intermediate stages of colorectal cancer diagnosis. Compared to static images, dynamic colonoscopy videos p…