3 citations · 6 across the 7 of their papers we have counts for
7 papers
Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning
Kazi Sajeed Mehrab, Hani Alomari, Najibul Haque Sarker +4
Multimodal large language models (MLLMs) ground whole objects well from free-form language queries, but they struggle when the query names a part rather than the object. We trace t…
NEST: Narrative Event Structures in Time for Long Video Understanding
Ali Asgarov, Kaushik Narasimhan, Najibul Haque Sarker +6
Recent progress in vision-language models has enabled processing of increasingly long video sequences, but handling extended token streams does not translate to understanding compl…
SONICS: Synthetic Or Not -- Identifying Counterfeit Songs
Md Awsafur Rahman, Zaber Ibn Abdul Hakim, Najibul Haque Sarker +2
The recent surge in AI-generated songs presents exciting possibilities and challenges. These innovations necessitate the ability to distinguish between human-composed and synthetic…
Leveraging Generative Language Models for Weakly Supervised Sentence Component Analysis in Video-Language Joint Learning
Zaber Ibn Abdul Hakim, Najibul Haque Sarker, Rahul Pratap Singh +3
A thorough comprehension of textual data is a fundamental element in multi-modal video analysis tasks. However, recent works have shown that the current models do not achieve a com…
Syn-Att: Synthetic Speech Attribution via Semi-Supervised Unknown Multi-Class Ensemble of CNNs
Md Awsafur Rahman, Bishmoy Paul, Najibul Haque Sarker +3
With the huge technological advances introduced by deep learning in audio & speech processing, many novel synthetic speech techniques achieved incredible realistic results. As thes…
Exploring Attention Mechanisms in Integration of Multi-Modal Information for Sign Language Recognition and Translation
Zaber Ibn Abdul Hakim, Rasman Mubtasim Swargo, Muhammad Abdullah Adnan
Understanding intricate and fast-paced movements of body parts is essential for the recognition and translation of sign language. The inclusion of additional information intended t…