#reward modeling

topicreward modeling

8 papers · 1 filter

cs.AI2026

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Qiushi Sun, Kanzhi Cheng, Yian Wang +20

The paper introduces OSReward, a benchmark for evaluating vision-language model judges that assess computer-using agent trajectories, and presents open reward models (OS‑Shepherd)…

cs.CL2026

FinanceHarness: Autonomous Financial Deep Research Framework

Yijia Xiao, Rujun Han, Yanfei Chen +8

The paper introduces FinanceHarness, a framework that uses large language models and autonomous agents to automate end‑to‑end financial deep research, and presents FinanceGym, a be…

cs.CL2026

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning

Jianze Wang, Kunwang Zheng, Ying Liu +5

The paper introduces SERPO, a test-time reinforcement learning approach that lets language models self‑improve during inference by jointly evolving response evidence, query‑specifi…

cs.CV2026

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation

Shihao Zhang, Yunzhi Li, Yuguang Yan +4

The paper introduces Shell-LCC, a method that treats the data manifold of high‑quality video training data as an implicit reward model, providing cheap, dense guidance for text‑to‑…

cs.SD2026

LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement

Chih-Ning Chen, Jen-Cheng Hou, Hsin-Min Wang +3

The paper introduces a reinforcement learning framework for audio‑visual speech enhancement that uses a large language model to generate natural‑language feedback, which is convert…

cs.RO2026

Deployable Human Preference Alignment in Robotics: Learning Representative Rewards from Diverse Human Preferences

Taehyung Kim, Gwangmo Lee, Minjun Chang +2

The paper proposes Preference-based REward Clustering (PREC), a method that groups users with similar preferences and learns a compact set of reward models from binary feedback to…