10 papers
Learning Quantifiable Visual Explanations Without Ground-Truth
Amritpal Singh, Andrey Barsky, Mohamed Ali Souibgui +2
Explainable AI (XAI) techniques are increasingly important for the validation and responsible use of modern deep learning models, but are difficult to evaluate due to the lack of g…
LayerBoost: Layer-Aware Attention Reduction for Efficient LLMs
Mohamed Ali Souibgui, Jan Fostier, Rodrigo AbadÃa-Heredia +3
Transformers are mostly relying on softmax attention, which introduces quadratic complexity with respect to sequence length and remains a major bottleneck for efficient inference.…
ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question Answering
Aymen Lassoued, Mohamed Ali Souibgui, Yousri Kessentini
Document Visual Question Answering (DocVQA) remains challenging for existing Vision-Language Models (VLMs), especially under complex reasoning and multi-step workflows. Current app…
NeurIPS 2023 Competition: Privacy Preserving Federated Learning Document VQA
Marlon Tobaben, Mohamed Ali Souibgui, Rubèn Tito +24
The Privacy Preserving Federated Learning Document VQA (PFL-DocVQA) competition challenged the community to develop provably private and communication-efficient solutions in a fede…
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
Mohamed Ali Souibgui, Changkyu Choi, Andrey Barsky +3
We propose DocVXQA, a novel framework for visually self-explainable document question answering. The framework is designed not only to produce accurate answers to questions but als…
One missing piece in Vision and Language: A Survey on Comics Understanding
Emanuele Vivoli, Mohamed Ali Souibgui, Andrey Barsky +3
Vision-language models have recently evolved into versatile systems capable of high performance across a range of tasks, such as document understanding, visual question answering,…