activity
20192026
most citedTrajectory Prediction for Autonomous Driving: Progress, Limitations, and Future Directions

32 citations · 93 across the 46 of their papers we have counts for

collaborators
Showing cs.CVShow all

39 papers · 1 filter

cs.CV2026

DERA: Detached Edge-Residual Adaptation for Prohibited item Detection

Yonathan Michael, Mohamad Alansari, Mohammed Bennamoun +3

Prohibited-item detection in X-ray imagery remains challenging due to object superposition, weak texture, and material clutter which obscure both semantic appearance and object bou…

cs.CV2026

CheXGround: Anatomical Region Tokens for Grounded Longitudinal Chest X-ray Interpretation

Adonay Demewez Gebremedhin, Wessam Shehieb, Sara Alansari +4

Recent radiology multi-modal language models have made substantial progress in chest X-ray report generation, visual question answering, and temporal reasoning. While longitudinal…

cs.CV2026

Polycepta: Object-Centric Appearance Estimation for Multi-Object Tracking

Mohamed Nagy, Naoufel Werghi, Jorge Dias +1

The tracking-by-detection paradigm in multi-object tracking (MOT) typically relies on static appearance descriptors to complement motion estimation. However, these descriptors are…

cs.CV2026

Falcon: Functional Assembly and Language for Compositional Reasoning in X-ray

Yonathan Michael, Mohamad Alansari, Natnael Takele +2

Conventional vision-language models are largely object-centric, focusing on detecting and describing individual entities. In safety-critical X-ray baggage screening, however, threa…

cs.CV2026

SENTRY: SAM2-Enhanced Neighbor-Aware and Temporally Reasoned Memory for Visual Tracking

Mohamad Alansari, Yonathan Michael, Hasan AlMarzouqi +3

We revisit the memory update mechanism in SAM2-based visual object tracking and identify confidence-only mask selection as the dominant cause of drift under occlusion, rapid motion…

cs.CV2026

SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs

Mohamad Alansari, Naufal Suryanto, Divya Velayudhan +3

Multimodal large language models (MLLMs) have advanced from image-level reasoning to pixel-level grounding, but extending these capabilities to videos remains challenging as models…