3 citations · 6 across the 8 of their papers we have counts for
6 papers · 1 filter
Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning
Kazi Sajeed Mehrab, Hani Alomari, Najibul Haque Sarker +4
Multimodal large language models (MLLMs) ground whole objects well from free-form language queries, but they struggle when the query names a part rather than the object. We trace t…
NEST: Narrative Event Structures in Time for Long Video Understanding
Ali Asgarov, Kaushik Narasimhan, Najibul Haque Sarker +6
Recent progress in vision-language models has enabled processing of increasingly long video sequences, but handling extended token streams does not translate to understanding compl…
LAMP: Learning Universal Adversarial Perturbations for Multi-Image Tasks via Pre-trained Models
Alvi Md Ishmam, Najibul Haque Sarker, Zaber Ibn Abdul Hakim +1
Multimodal Large Language Models (MLLMs) have achieved remarkable performance across vision-language tasks. Recent advancements allow these models to process multiple images as inp…
An Optimized YOLOv5 Based Approach For Real-time Vehicle Detection At Road Intersections Using Fisheye Cameras
Md. Jahin Alam, Muhammad Zubair Hasan, Md Maisoon Rahman +8
Real time vehicle detection is a challenging task for urban traffic surveillance. Increase in urbanization leads to increase in accidents and traffic congestion in junction areas r…
ENTER: Event Based Interpretable Reasoning for VideoQA
Hammad Ayyubi, Junzhang Liu, Ali Asgarov +10
In this paper, we present ENTER, an interpretable Video Question Answering (VideoQA) system based on event graphs. Event graphs convert videos into graphical representations, where…
ArtiFact: A Large-Scale Dataset with Artificial and Factual Images for Generalizable and Robust Synthetic Image Detection
Md Awsafur Rahman, Bishmoy Paul, Najibul Haque Sarker +2
Synthetic image generation has opened up new opportunities but has also created threats in regard to privacy, authenticity, and security. Detecting fake images is of paramount impo…