1 citations · 4 across the 17 of their papers we have counts for
1 paper · 1 filter
Yin Li, Ziyang Hu, Zhiyu Guo +7
Existing multimodal RAG methods often flatten structured documents into isolated text and image units, weakening the source organization and local text-image logic needed for faith…