2 papers
cs.CV2026
Faithful Grounded Visual Reasoning via Learned Proxy-Tokens
Tom Hodemon, Mohamed Chaouch, Aboubacar Tuo +1
Multimodal Large Language Models (MLLMs) have achieved remarkable success in Visual Question Answering (VQA), yet their "black-box" nature hinders deployment in critical domains. G…
cs.CV2026
Benchmarking Adversarial Robustness and Adversarial Training Strategies for Object Detection
Alexis Winter, Jean-Vincent Martini, Romaric Audigier +2
Object detection models are critical components of automated systems, such as autonomous vehicles and perception-based robots, but their sensitivity to adversarial attacks poses a…