2 papers
cs.LG2026
The Weakest Link Tells It All: Outcome-Supervised Process Reward Modeling via Learnable Credit Assignment
Tianyu Jia, Yue Fang, Hongxin Ding +6
Process reward models (PRMs) enhance the reasoning capabilities of large language models (LLMs) by providing fine-grained feedback, yet training PRMs typically requires expensive s…
cs.LG2025
Stackelberg Self-Annotation: A Robust Approach to Data-Efficient LLM Alignment
Xu Chu, Zhixin Zhang, Tianyu Jia +1
Aligning large language models (LLMs) with human preferences typically demands vast amounts of meticulously curated data, which is both expensive and prone to labeling noise. We pr…