5 citations · 12 across the 25 of their papers we have counts for
22 papers · 1 filter
AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images
Abderrahmene Boudiaf, Mohamad Alanssari, Irfan Hussain +1
Agricultural image understanding requires fine-grained recognition of plant diseases, pests, crop structures, and botanical species under complex real-world conditions. Despite rec…
CheXGround: Anatomical Region Tokens for Grounded Longitudinal Chest X-ray Interpretation
Adonay Demewez Gebremedhin, Wessam Shehieb, Sara Alansari +4
Recent radiology multi-modal language models have made substantial progress in chest X-ray report generation, visual question answering, and temporal reasoning. While longitudinal…
From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology
Basit Alawode, Moshira Ali Abdalla, Dwarikanath Mahapatra +2
Vision Transformers (ViTs) and their hierarchical variants have achieved strong performance in Computational Pathology (CPath). However, most are pre-trained on single-resolution W…
SENTRY: SAM2-Enhanced Neighbor-Aware and Temporally Reasoned Memory for Visual Tracking
Mohamad Alansari, Yonathan Michael, Hasan AlMarzouqi +3
We revisit the memory update mechanism in SAM2-based visual object tracking and identify confidence-only mask selection as the dominant cause of drift under occlusion, rapid motion…
SegRAG: Training-Free Retrieval-Augmented Semantic Segmentation
Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed
Open-vocabulary segmentation models such as SAM3 perform well across broad categories via text prompting, yet degrade when target classes are visually underrepresented in pretraini…
MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image Understanding
Basit Alawode, Arif Mahmood, Muaz Khalifa Al-Radi +6
Whole Slide Images (WSIs) exhibit hierarchical structure, where diagnostic information emerges from cellular morphology, regional tissue organization, and global context. Existing…