3 papers
cs.CV2025
Representation Learning with Adaptive Superpixel Coding
Mahmoud Khalil, Ahmad Khalil, Alioune Ngom
Deep learning vision models are typically tailored for specific modalities and often rely on domain-specific assumptions, such as the grid structures used by nearly all existing vi…
cs.CV2025
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task
Ahmad Khalil, Mahmoud Khalil, Alioune Ngom
In this paper, we introduce ResNetVLLM (ResNet Vision LLM), a novel cross-modal framework for zero-shot video understanding that integrates a ResNet-based visual encoder with a Lar…
cs.CV2025
ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations
Ahmad Khalil, Mahmoud Khalil, Alioune Ngom
Large Language Models (LLMs) have transformed natural language processing (NLP) tasks, but they suffer from hallucination, generating plausible yet factually incorrect content. Thi…