Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Faithful Grounded Visual Reasoning via Learned Proxy-Tokens
Tom Hodemon, Mohamed Chaouch, Aboubacar Tuo +1
Multimodal Large Language Models (MLLMs) have achieved remarkable success in Visual Question Answering (VQA), yet their "black-box" nature hinders deployment in critical domains. G…
cs.CV2026
Improving Controllable Generation: Faster Training and Better Performance via -Supervision
Amadou S. Sangare, Adrien Maglo, Mohamed Chaouch +1
Text-to-Image (T2I) diffusion/flow models have recently achieved remarkable progress in visual fidelity and text alignment. However, they remain limited when users need to precisel…
cs.CV2024
3D-COCO: extension of MS-COCO dataset for image detection and 3D reconstruction modules
Maxence Bideaux, Alice Phe, Mohamed Chaouch +2
We introduce 3D-COCO, an extension of the original MS-COCO dataset providing 3D models and 2D-3D alignment annotations. 3D-COCO was designed to achieve computer vision tasks such a…