4 papers · 1 filter
Representation Learning with Adaptive Superpixel Coding
Mahmoud Khalil, Ahmad Khalil, Alioune Ngom
Deep learning vision models are typically tailored for specific modalities and often rely on domain-specific assumptions, such as the grid structures used by nearly all existing vi…
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task
Ahmad Khalil, Mahmoud Khalil, Alioune Ngom
In this paper, we introduce ResNetVLLM (ResNet Vision LLM), a novel cross-modal framework for zero-shot video understanding that integrates a ResNet-based visual encoder with a Lar…
ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations
Ahmad Khalil, Mahmoud Khalil, Alioune Ngom
Large Language Models (LLMs) have transformed natural language processing (NLP) tasks, but they suffer from hallucination, generating plausible yet factually incorrect content. Thi…
A Comprehensive Study of Vision Transformers in Image Classification Tasks
Mahmoud Khalil, Ahmad Khalil, Alioune Ngom
Image Classification is a fundamental task in the field of computer vision that frequently serves as a benchmark for gauging advancements in Computer Vision. Over the past few year…