#reward modeling
8 papers · 1 filter
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Qiushi Sun, Kanzhi Cheng, Yian Wang +20
The paper introduces OSReward, a benchmark for evaluating vision-language model judges that assess computer-using agent trajectories, and presents open reward models (OS‑Shepherd)…
FinanceHarness: Autonomous Financial Deep Research Framework
Yijia Xiao, Rujun Han, Yanfei Chen +8
The paper introduces FinanceHarness, a framework that uses large language models and autonomous agents to automate end‑to‑end financial deep research, and presents FinanceGym, a be…
SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning
Jianze Wang, Kunwang Zheng, Ying Liu +5
The paper introduces SERPO, a test-time reinforcement learning approach that lets language models self‑improve during inference by jointly evolving response evidence, query‑specifi…
Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation
Shihao Zhang, Yunzhi Li, Yuguang Yan +4
The paper introduces Shell-LCC, a method that treats the data manifold of high‑quality video training data as an implicit reward model, providing cheap, dense guidance for text‑to‑…
LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
Chih-Ning Chen, Jen-Cheng Hou, Hsin-Min Wang +3
The paper introduces a reinforcement learning framework for audio‑visual speech enhancement that uses a large language model to generate natural‑language feedback, which is convert…
Deployable Human Preference Alignment in Robotics: Learning Representative Rewards from Diverse Human Preferences
Taehyung Kim, Gwangmo Lee, Minjun Chang +2
The paper proposes Preference-based REward Clustering (PREC), a method that groups users with similar preferences and learns a compact set of reward models from binary feedback to…