3 papers
cs.CV2025
A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
Phu-Vinh Nguyen, Tan-Hanh Pham, Chris Ngo +1
The development of many vision models mainly focuses on improving their performance using metrics such as accuracy, IoU, and mAP, with less attention to explainability due to the c…
cs.CV2025
IQBench: How "Smart'' Are Vision-Language Models? A Study with Human IQ Tests
Tan-Hanh Pham, Phu-Vinh Nguyen, Dang The Hung +5
Although large Vision-Language Models (VLMs) have demonstrated remarkable performance in a wide range of multimodal tasks, their true reasoning capabilities on human IQ tests remai…
cs.CV2024
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization
Tan-Hanh Pham, Hoang-Nam Le, Phu-Vinh Nguyen +2
Visual Language Models have demonstrated remarkable capabilities across tasks, including visual question answering and image captioning. However, most models rely on text-based ins…