3 papers
cs.CV2025
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
Brigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh +5
Vision-Language Models (VLMs) have achieved remarkable progress across tasks such as visual question answering and image captioning. Yet, the extent to which these models perform v…
cs.CV2025
Vision-Language Models Can't See the Obvious
Yasser Dahou, Ngoc Dung Huynh, Phuc H. Le-Khac +3
We present Saliency Benchmark (SalBench), a novel benchmark designed to assess the capability of Large Vision-Language Models (LVLM) in detecting visually salient features that are…
cs.AI2025
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models
Mouadh Yagoubi, Yasser Dahou, Billel Mokeddem +12
Existing benchmarks have proven effective for assessing the performance of fully trained large language models. However, we find striking differences in the early training stages o…