15 citations · 16 across the 6 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Toward Native Multimodal Modeling: A Roadmap
Siyu An, Junru Lu, Junnan Dong +18
Multimodal modeling represents a vital step from modality-agnostic reasoning toward world modeling. While early approaches predominantly rely on late-fusion that assembles encoders…
cs.CV2024
Modality-Aware Integration with Large Language Models for Knowledge-based Visual Question Answering
Junnan Dong, Qinggang Zhang, Huachi Zhou +3
Knowledge-based visual question answering (KVQA) has been extensively studied to answer visual questions with external knowledge, e.g., knowledge graphs (KGs). While several attemp…