82 citations · 289 across the 39 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
GraphVLM: Benchmarking Vision Language Models for Multimodal Graph Learning
Jiajin Liu, Dongzhe Fan, Chuanhao Ji +2
Vision-Language Models (VLMs) have demonstrated remarkable capabilities in aligning and understanding multimodal signals, yet their potential to reason over structured data, where…
cs.CV2024
Modality-Aware Integration with Large Language Models for Knowledge-based Visual Question Answering
Junnan Dong, Qinggang Zhang, Huachi Zhou +3
Knowledge-based visual question answering (KVQA) has been extensively studied to answer visual questions with external knowledge, e.g., knowledge graphs (KGs). While several attemp…