7 papers
Spectral Souping: A Unified Framework for Online Preference Alignment
Yinlam Chow, Guy Tennenholtz, Ted Yun +4
Reinforcement Learning from Human Feedback (RLHF) effectively aligns Large Language Models (LLMs) with aggregate human preferences but often fails to address the diverse and confli…
Controllable User Simulation
Guy Tennenholtz, Ofer Meshi, Amir Globerson +3
Using offline datasets to evaluate conversational agents often fails to cover rare scenarios or to support testing new policies. This has motivated the use of controllable user sim…
Diffusion Controller: Framework, Algorithms and Parameterization
Tong Yang, Moonkyung Ryu, Chih-Wei Hsu +4
Controllable diffusion generation often relies on various heuristics that are seemingly disconnected without a unified understanding. We bridge this gap with Diffusion Controller (…
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
Yinlam Chow, Guy Tennenholtz, Izzeddin Gur +7
Recent studies have indicated that effectively utilizing inference-time compute is crucial for attaining better performance from large language models (LLMs). In this work, we prop…
Asking Clarifying Questions for Preference Elicitation With Large Language Models
Ali Montazeralghaem, Guy Tennenholtz, Craig Boutilier +1
Large Language Models (LLMs) have made it possible for recommendation systems to interact with users in open-ended conversational interfaces. In order to personalize LLM responses,…
Descriptive History Representations: Learning Representations by Answering Questions
Guy Tennenholtz, Jihwan Jeong, Chih-Wei Hsu +2
Effective decision making in partially observable environments requires compressing long interaction histories into informative representations. We introduce Descriptive History Re…