2 papers
cs.CV2026
Faithful Grounded Visual Reasoning via Learned Proxy-Tokens
Tom Hodemon, Mohamed Chaouch, Aboubacar Tuo +1
Multimodal Large Language Models (MLLMs) have achieved remarkable success in Visual Question Answering (VQA), yet their "black-box" nature hinders deployment in critical domains. G…
cs.CV2026
Improving Controllable Generation: Faster Training and Better Performance via -Supervision
Amadou S. Sangare, Adrien Maglo, Mohamed Chaouch +1
Text-to-Image (T2I) diffusion/flow models have recently achieved remarkable progress in visual fidelity and text alignment. However, they remain limited when users need to precisel…