4 papers · 1 filter
Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
Ali Rasekh, Erfan Bagheri Soula, Omid Daliran +2
Despite significant advances in Multimodal Large Language Models (MLLMs), understanding complex temporal dynamics in videos remains a major challenge. Our experiments show that cur…
Multi-Rationale Explainable Object Recognition via Contrastive Conditional Inference
Ali Rasekh, Sepehr Kazemi Ranjbar, Simon Gottschalk
Explainable object recognition using vision-language models such as CLIP involves predicting accurate category labels supported by rationales that justify the decision-making proce…
Aligning Visual Contrastive learning models via Preference Optimization
Amirabbas Afzali, Borna Khodabandeh, Ali Rasekh +3
Contrastive learning models have demonstrated impressive abilities to capture semantic similarities by aligning representations in the embedding space. However, their performance c…
ECOR: Explainable CLIP for Object Recognition
Ali Rasekh, Sepehr Kazemi Ranjbar, Milad Heidari +1
Large Vision Language Models (VLMs), such as CLIP, have significantly contributed to various computer vision tasks, including object recognition and object detection. Their open vo…