3 papers
cs.AI2026
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF
Wei Chen, Guanghui Zhu, Yafei Li +2
Reinforcement learning from human feedback (RLHF) with preference-based reward models often exhibits unstable training dynamics. A key contributing factor is that standard RLHF rel…
cs.MM2025
MHier-RAG: Multi-Modal RAG for Visual-Rich Document Question-Answering via Hierarchical and Multi-Granularity Reasoning
Ziyu Gong, Chengcheng Mai, Yihua Huang
The multi-modal long-context document question-answering task aims to locate and integrate multi-modal evidences (such as texts, tables, charts, images, and layouts) distributed ac…
cs.CL2025
KnowRA: Knowledge Retrieval Augmented Method for Document-level Relation Extraction with Comprehensive Reasoning Abilities
Chengcheng Mai, Yuxiang Wang, Ziyu Gong +2
Document-level relation extraction (Doc-RE) aims to extract relations between entities across multiple sentences. Therefore, Doc-RE requires more comprehensive reasoning abilities…