#visual question answering

try —

9 papers match

cs.AI2026

Distilling Answer Set Programming Theories from Large Language Models

Nelson Higuera Ruiz, Markus Hofmarcher, Claudiu Leoveanu-Condrei

The paper investigates whether large language models can automatically generate correct Answer Set Programming theories for visual question answering tasks within a one‑hour time l…

#answer set programming#neurosymbolic learning#large language models#visual question answering
cs.AI2026

PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?

Zongyi Chen, Yu Liang, Jie Lin +1

The paper presents PathVU, a benchmark that tests multimodal large language models on fine-grained, multiscale visual understanding of pathology images using region- and slide-leve…

#multimodal large language models#pathology imaging#visual question answering#multiscale understanding
cs.LG2026

LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger

Enjun Du, Hange Zhou, Chenxu Du +4

The paper introduces LedgerMind, a framework that records and constrains the evidence used by multimodal agents during visual question answering, ensuring that each reasoning step…

#multimodal reasoning#visual question answering#provenance tracking#grounded reasoning
cs.CV2026

ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures

Fahad Ahmed, Sören Auer, Jennifer D'Souza

The paper presents the ICDAR 2026 competition and the Sci-ImageMiner benchmark for extracting and reasoning over information in scientific figures related to atomic layer depositio…

#scientific figure understanding#multimodal learning#information extraction#visual question answering
cs.CV2026

Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs

Itbaan Safwan, Ramail Khan, Muhammad Annas Shaikh +1

The paper introduces a multi‑task fine‑tuning approach for small vision‑language models to improve visual question answering on GI endoscopy images, adding grounding and descriptio…

#visual question answering#gastrointestinal endoscopy#multitask learning#vision-language models
cs.LG2026

Regularizing modality contribution drift in multimodal continual learning

Zhen Zhang, Jielei Chu, Bin Liu +1

The paper identifies a decision-level shift called Modality Contribution Drift in multimodal continual learning and introduces a regularization method (CMCDR) that preserves modali…

#multimodal learning#continual learning#modality contribution drift#regularization