#preference learning
9 papers match
AutoPref: Automatic Discovery of Task-Specific Preference Objectives for Neural Combinatorial Optimization
Shengda Gu, Kai Li, Xinyi Ke +3
AutoPref uses a large language model to automatically discover and compose pairwise loss and weighting programs that define preference objectives for neural combinatorial optimizat…
Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence
Haocheng Yang, Licheng Pan, Xiaoxi Li +5
The paper proposes a query‑only method that automatically creates and validates fine‑grained rubrics for evaluating large language models by using synthetic rubric‑conditioned resp…
Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs
Zhixiao Zheng, Zheren Fu, Zhiyuan Yao +3
The paper introduces Groc-PO, a preference‑optimization framework that provides stage‑specific supervision for object grounding, contextual grounding, and grounded reasoning in mul…
RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences
Logan Mondal Bhamidipaty, Mykel Kochenderfer, Subramanian Ramamoorthy
The paper introduces RENEW, a method that uses human preferences over imagined rollouts to correct model exploitation in offline model-based reinforcement learning, focusing fine‑t…
Deployable Human Preference Alignment in Robotics: Learning Representative Rewards from Diverse Human Preferences
Taehyung Kim, Gwangmo Lee, Minjun Chang +2
The paper proposes Preference-based REward Clustering (PREC), a method that groups users with similar preferences and learns a compact set of reward models from binary feedback to…
Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models
Ilias Kazantzidis, Timothy J. Norman, Yali Du +1
The paper introduces DROPJ, a human‑in‑the‑loop approach that uses preferences and justifications collected from users interacting with a learned world model to train a reward mode…
Meta-Learning Preferences for Multilingual LLM Alignment
Jiaying Lin, Seongho Son, Nam Phuong Tran +3
The paper introduces a meta-learning method that uses preference data from high-resource languages to quickly adapt large language models to low-resource languages with very few hu…
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability
Simone Drago, Marco Mussi, Leonardo Bianconi +1
The paper extends preference‑based reinforcement learning by allowing human experts to label trajectory pairs as incomparable, and introduces a Bradley‑Terry‑inspired rationality m…
Freeform Preference Learning for Robotic Manipulation
Marcel Torne, Anubha Mahajan, Abhijnya Bhat +1
The paper introduces Freeform Preference Learning, a method that lets humans give natural-language preference criteria for robot trajectories, enabling robots to learn multi-dimens…
One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.