3 papers
cs.CV2026
Modern Backbones Improve Multi-task DETR for Mammography Classification and Lesion Localization
Dinh Tan Nguyen, Quang-Hien Kha, Le-Hoang Nguyen +8
Joint exam-level prediction and candidate-region localization may improve the usefulness of AI support in mammography. We study this setting using a multi-task DETR framework, wher…
cs.CV2025
GMAT: Grounded Multi-Agent Clinical Description Generation for Text Encoder in Vision-Language MIL for Whole Slide Image Classification
Ngoc Bui Lam Quang, Nam Le Nguyen Binh, Thanh-Huy Nguyen +3
Multiple Instance Learning (MIL) is the leading approach for whole slide image (WSI) classification, enabling efficient analysis of gigapixel pathology slides. Recent work has intr…
cs.CV2025
Describe Anything Model for Visual Question Answering on Text-rich Images
Yen-Linh Vu, Dinh-Thang Duong, Truong-Binh Duong +8
Recent progress has been made in region-aware vision-language modeling, particularly with the emergence of the Describe Anything Model (DAM). DAM is capable of generating detailed…