3 papers
cs.CV2025
PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to Histopathology
Fatemeh Ghezloo, Mehmet Saygin Seyfioglu, Rustin Soraki +6
Diagnosing diseases through histopathology whole slide images (WSIs) is fundamental in modern pathology but is challenged by the gigapixel scale and complexity of WSIs. Trained his…
cs.CV2025
MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives
Wisdom O. Ikezogwo, Kevin Zhang, Mehmet Saygin Seyfioglu +3
Multi-modal models are data hungry. While datasets with natural images are abundant, medical image datasets can not afford the same luxury. To enable representation learning for me…
cs.CV2024
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
Mahtab Bigverdi, Zelun Luo, Cheng-Yu Hsieh +4
Multimodal language models (MLMs) still face challenges in fundamental visual perception tasks where specialized models excel. Tasks requiring reasoning about 3D structures benefit…