3 papers
cs.IR2025
Contextual Relevance and Adaptive Sampling for LLM-Based Document Reranking
Jerry Huang, Siddarth Madala, Cheng Niu +2
Reranking algorithms have made progress in improving document retrieval quality by efficiently aggregating relevance judgments generated by large language models (LLMs). However, i…
cs.CL2025
DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning
Yuanhao Wu, Juntong Song, Hanning Zhang +2
In this paper, we propose DuaShepherd, a novel reward modeling framework that integrates two complementary reward signals, correctness and potential, to enhance the mathematical re…
cs.CL2025
OpenGenAlign: A Preference Dataset and Benchmark for Trustworthy Reward Modeling in Open-Ended, Long-Context Generation
Hanning Zhang, Juntong Song, Juno Zhu +3
Reward Modeling is critical in evaluating and improving the generation of Large Language Models (LLMs). While numerous recent works have shown its feasibility in improving safety,…