From the 1 of 21 linked papers with an AI index.
21 papers
PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models
Vignesh Kothapalli, Rishabh Ranjan, Valter Hudovernik +4
The paper presents PLUREL, a lightweight framework for generating synthetic multi-table relational databases, enabling the study of scaling laws in relational foundation models and…
Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds
Edward Chen, Sang T. Truong, Natalie Dullerud +2
High-stakes decision-making involves navigating multiple competing objectives with expensive evaluations. For instance, in brachytherapy, clinicians must balance maximizing tumor c…
Beyond expert users: agents should help users construct preferences, not just elicit them
Irena Saracay, Ludwig Schmidt, Carlos Guestrin
Agents typically assume an expert user -- one with well-formed preferences about what they want -- and default to clarifying questions whenever the task is underspecified. We argue…
AI Assistance for Human Review of Default Judgments
Theodora Worledge, Othman Bensouda Koraichi, Daniel Bernal +4
Overwhelmed courts in the United States review millions of default judgments each year. Unfortunately, such manual reviews are time-consuming and prone to error. In an audit of 188…
Discovering Implicit Large Language Model Alignment Objectives
Edward Chen, Sanmi Koyejo, Carlos Guestrin
Large language model (LLM) alignment relies on complex reward signals that often obscure the specific behaviors being incentivized, creating critical risks of misalignment and rewa…
Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning
Qinan Yu, Alexa Tartaglini, Peter Hase +2
Reinforcement Learning from Verifiable Rewards (RLVR) on chain-of-thought reasoning has become a standard part of language model post-training recipes. A common assumption is that…