1 paper · 1 filter
Aymen Lassoued, Mohamed Ali Souibgui, Yousri Kessentini
Document Visual Question Answering (DocVQA) remains challenging for existing Vision-Language Models (VLMs), especially under complex reasoning and multi-step workflows. Current app…