2 papers
cs.CV2026
MaS-VQA: A Mask-and-Select Framework for Knowledge-Based Visual Question Answering
Xianwei Mao, Kai Ye, Sheng Zhou +4
Knowledge-based Visual Question Answering (KB-VQA) requires models to answer questions by integrating visual information with external knowledge. However, retrieved knowledge is of…
cs.CV2025
Multimodal Machine Translation with Visual Scene Graph Pruning
Chenyu Lu, Shiliang Sun, Jing Zhao +3
Multimodal machine translation (MMT) seeks to address the challenges posed by linguistic polysemy and ambiguity in translation tasks by incorporating visual information. A key bott…