2 citations · 2 across the 2 of their papers we have counts for
4 papers
Learning Explainable Dense Reward Shapes via Bayesian Optimization
Ryan Koo, Ian Yang, Vipul Raheja +3
Current reinforcement learning from human feedback (RLHF) pipelines for large language model (LLM) alignment typically assign scalar rewards to sequences, using the final token as…
Dynamic Multi-Reward Weighting for Multi-Style Controllable Generation
Karin de Langis, Ryan Koo, Dongyeop Kang
Textual style expresses a diverse set of information, including interpersonal dynamics (e.g., formality) and the author's emotions or attitudes (e.g., disgust). An open question is…
Under the Surface: Tracking the Artifactuality of LLM-Generated Data
Debarati Das, Karin De Langis, Anna Martin-Boyle +14
This work delves into the expanding role of large language models (LLMs) in generating artificial data. LLMs are increasingly employed to create a variety of outputs, including ann…
Benchmarking Cognitive Biases in Large Language Models as Evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja +3
Large Language Models are cognitively biased judges. Large Language Models (LLMs) have recently been shown to be effective as automatic evaluators with simple prompting and in-cont…