2 papers
cs.LG2025
Learning Explainable Dense Reward Shapes via Bayesian Optimization
Ryan Koo, Ian Yang, Vipul Raheja +3
Current reinforcement learning from human feedback (RLHF) pipelines for large language model (LLM) alignment typically assign scalar rewards to sequences, using the final token as…
cs.CL2024
Dynamic Multi-Reward Weighting for Multi-Style Controllable Generation
Karin de Langis, Ryan Koo, Dongyeop Kang
Textual style expresses a diverse set of information, including interpersonal dynamics (e.g., formality) and the author's emotions or attitudes (e.g., disgust). An open question is…