7 papers
GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model
Peter Bohm, Saimunur Rahman, Abdelwahed Khamis +3
Visual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated visual encoder, a bespoke action head, and training on thousands of ho…
SteerSeg: Attention Steering for Reasoning Video Segmentation
Ali Cheraghian, Hamidreza Dastmalchi, Abdelwahed Khamis +3
Video reasoning segmentation requires localizing objects across video frames from natural language expressions, often involving spatial reasoning and implicit references. Recent ap…
TPCL: Task Progressive Curriculum Learning for Robust Visual Question Answering
Ahmed Akl, Abdelwahed Khamis, Zhe Wang +3
Visual Question Answering (VQA) systems are notoriously brittle under distribution shifts and data scarcity. While previous solutions-such as ensemble methods and data augmentation…
Component-Aware Sketch-to-Image Generation Using Self-Attention Encoding and Coordinate-Preserving Fusion
Ali Zia, Muhammad Umer Ramzan, Usman Ali +3
Translating freehand sketches into photorealistic images remains a fundamental challenge in image synthesis, particularly due to the abstract, sparse, and stylistically diverse nat…
HIME: Mitigating Object Hallucinations in LVLMs via Hallucination Insensitivity Model Editing
Ahmed Akl, Abdelwahed Khamis, Ali Cheraghian +3
Large Vision-Language Models (LVLMs) have demonstrated impressive multimodal understanding capabilities, yet they remain prone to object hallucination, where models describe non-ex…
NeuralPrefix: A Zero-shot Sensory Data Imputation Plugin
Abdelwahed Khamis, Sara Khalifa
Real-world sensing challenges such as sensor failures, communication issues, and power constraints lead to data intermittency. An issue that is known to undermine the traditional c…