#visual question answering
9 papers match
Distilling Answer Set Programming Theories from Large Language Models
Nelson Higuera Ruiz, Markus Hofmarcher, Claudiu Leoveanu-Condrei
The paper investigates whether large language models can automatically generate correct Answer Set Programming theories for visual question answering tasks within a one‑hour time l…
PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?
Zongyi Chen, Yu Liang, Jie Lin +1
The paper presents PathVU, a benchmark that tests multimodal large language models on fine-grained, multiscale visual understanding of pathology images using region- and slide-leve…
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
Enjun Du, Hange Zhou, Chenxu Du +4
The paper introduces LedgerMind, a framework that records and constrains the evidence used by multimodal agents during visual question answering, ensuring that each reasoning step…
ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures
Fahad Ahmed, Sören Auer, Jennifer D'Souza
The paper presents the ICDAR 2026 competition and the Sci-ImageMiner benchmark for extracting and reasoning over information in scientific figures related to atomic layer depositio…
Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs
Itbaan Safwan, Ramail Khan, Muhammad Annas Shaikh +1
The paper introduces a multi‑task fine‑tuning approach for small vision‑language models to improve visual question answering on GI endoscopy images, adding grounding and descriptio…
Regularizing modality contribution drift in multimodal continual learning
Zhen Zhang, Jielei Chu, Bin Liu +1
The paper identifies a decision-level shift called Modality Contribution Drift in multimodal continual learning and introduces a regularization method (CMCDR) that preserves modali…