9 papers
Learning Latent Reasoning Traces for Scalar Reward Models End-to-End
Sanwoo Lee, Clive Bai, Hsiu-Yuan Huang +3
Reward models (RMs) are central to aligning large language models with human preferences via reinforcement learning. Although traditional scalar RMs enable efficient and probabilis…
RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data
Hsiu-Yuan Huang, Weijie Liu, Chenming Tang +5
The proliferation of Reinforcement Learning from Verifiable Rewards (RLVR) datasets has exacerbated provenance collapse due to unclear lineage among existing datasets. To bridge th…
Trait-Aware Policy Optimization for Autoregressive Multi-Trait Essay Scoring
Zhengyang Wang, Sanwoo Lee, Jiaxin Wang +3
Multi-trait essay scoring aims to provide fine-grained evaluation of writing quality across multiple dimensions. However, how to effectively post-train autoregressive scoring model…
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
Kun Liang, Clive Bai, Xin Xu +5
Recent Large Reasoning Models (LRMs) achieve strong performance by leveraging long-form Chain-of-Thought (CoT) reasoning, but uniformly applying overlong reasoning at inference tim…
Composable Cross-prompt Essay Scoring by Merging Models
Sanwoo Lee, Kun Liang, Yunfang Wu
Recent advances in cross-prompt automated essay scoring (AES) typically train models jointly on all source prompts, often requiring additional access to unlabeled target prompt ess…
Dynamic Fisher-weighted Model Merging via Bayesian Optimization
Sanwoo Lee, Jiahao Liu, Qifan Wang +3
The fine-tuning of pre-trained language models has resulted in the widespread availability of task-specific models. Model merging offers an efficient way to create multi-task model…