1 paper
Yikai Gao, Ding Xia, Xi Yang
In domain-specific multimodal long documents, images and text jointly convey complex knowledge that cannot be fully captured by plain text alone. However, existing paradigms like D…