1 paper
Yin Li, Ziyang Hu, Zhiyu Guo +7
Existing multimodal RAG methods often flatten structured documents into isolated text and image units, weakening the source organization and local text-image logic needed for faith…