2 papers
cs.IR2025
HKRAG: Holistic Knowledge Retrieval-Augmented Generation Over Visually-Rich Documents
Anyang Tong, Xiang Niu, ZhiPing Liu +4
Existing multimodal Retrieval-Augmented Generation (RAG) methods for visually rich documents (VRD) are often biased towards retrieving salient knowledge(e.g., prominent text and vi…
cs.AI2025
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models
Xiangxiang Zhang, Jingxuan Wei, Donghong Zhong +31
Existing Vision-Language Models often struggle with complex, multi-question reasoning tasks where partial correctness is crucial for effective learning. Traditional reward mechanis…