4 papers
RAG-Reflect: Agentic Retrieval-Augmented Generation with Reflections for Comment-Driven Code Maintenance on Stack Overflow
Mehedi Hasan Shanto, Muhammad Asaduzzaman, Alioune Ngom
User comments on online programming platforms such as Stack Overflow play a vital role in maintaining the correctness and relevance of shared code examples. However, the majority o…
Representation Learning with Adaptive Superpixel Coding
Mahmoud Khalil, Ahmad Khalil, Alioune Ngom
Deep learning vision models are typically tailored for specific modalities and often rely on domain-specific assumptions, such as the grid structures used by nearly all existing vi…
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task
Ahmad Khalil, Mahmoud Khalil, Alioune Ngom
In this paper, we introduce ResNetVLLM (ResNet Vision LLM), a novel cross-modal framework for zero-shot video understanding that integrates a ResNet-based visual encoder with a Lar…
ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations
Ahmad Khalil, Mahmoud Khalil, Alioune Ngom
Large Language Models (LLMs) have transformed natural language processing (NLP) tasks, but they suffer from hallucination, generating plausible yet factually incorrect content. Thi…