4 papers
Uncertainty-Aware Decision Making in Multimodal Large Language Models
Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed
Multimodal large language models (MLLMs) increasingly answer questions whose correctness depends on visual, textual, temporal, acoustic, document, chart, or embodied evidence. Thei…
SegRAG: Training-Free Retrieval-Augmented Semantic Segmentation
Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed
Open-vocabulary segmentation models such as SAM3 perform well across broad categories via text prompting, yet degrade when target classes are visually underrepresented in pretraini…
Integrating Features for Recognizing Human Activities through Optimized Parameters in Graph Convolutional Networks and Transformer Architectures
Mohammad Belal, Taimur Hassan, Abdelfatah Hassan +3
Human activity recognition is a major field of study that employs computer vision, machine vision, and deep learning techniques to categorize human actions. The field of deep learn…
Feature Fusion for Human Activity Recognition using Parameter-Optimized Multi-Stage Graph Convolutional Network and Transformer Models
Mohammad Belal, Taimur Hassan, Abdelfatah Ahmed +3
Human activity recognition (HAR) is a crucial area of research that involves understanding human movements using computer and machine vision technology. Deep learning has emerged a…