32 citations · 93 across the 46 of their papers we have counts for
39 papers · 1 filter
DERA: Detached Edge-Residual Adaptation for Prohibited item Detection
Yonathan Michael, Mohamad Alansari, Mohammed Bennamoun +3
Prohibited-item detection in X-ray imagery remains challenging due to object superposition, weak texture, and material clutter which obscure both semantic appearance and object bou…
CheXGround: Anatomical Region Tokens for Grounded Longitudinal Chest X-ray Interpretation
Adonay Demewez Gebremedhin, Wessam Shehieb, Sara Alansari +4
Recent radiology multi-modal language models have made substantial progress in chest X-ray report generation, visual question answering, and temporal reasoning. While longitudinal…
Polycepta: Object-Centric Appearance Estimation for Multi-Object Tracking
Mohamed Nagy, Naoufel Werghi, Jorge Dias +1
The tracking-by-detection paradigm in multi-object tracking (MOT) typically relies on static appearance descriptors to complement motion estimation. However, these descriptors are…
Falcon: Functional Assembly and Language for Compositional Reasoning in X-ray
Yonathan Michael, Mohamad Alansari, Natnael Takele +2
Conventional vision-language models are largely object-centric, focusing on detecting and describing individual entities. In safety-critical X-ray baggage screening, however, threa…
SENTRY: SAM2-Enhanced Neighbor-Aware and Temporally Reasoned Memory for Visual Tracking
Mohamad Alansari, Yonathan Michael, Hasan AlMarzouqi +3
We revisit the memory update mechanism in SAM2-based visual object tracking and identify confidence-only mask selection as the dominant cause of drift under occlusion, rapid motion…
SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs
Mohamad Alansari, Naufal Suryanto, Divya Velayudhan +3
Multimodal large language models (MLLMs) have advanced from image-level reasoning to pixel-level grounding, but extending these capabilities to videos remains challenging as models…