From the 1 of 17 linked papers with an AI index.
17 papers
Efficient Online Proportional Sampling with Applications to Smoothed Online Learning
Amirmahdi Mirfakhar, Maria-Florina Balcan, Hedyeh Beyhaghi
The paper proposes a data structure for efficiently performing online proportional sampling over high‑dimensional, piecewise‑structured domains under a σ‑smoothed adversary, and us…
Reward Learning from Best-of- Preference Data: Targets, Tradeoffs, and Design Principles
Rattana Pukdee, Maria-Florina Balcan, Pradeep Ravikumar
Best-of- sampling is widely used to construct pairwise preference data: candidates are drawn from a base distribution, and the best is paired with a rejected response. Despi…
What Does Preference Learning Recover from Pairwise Comparison Data?
Rattana Pukdee, Maria-Florina Balcan, Pradeep Ravikumar
Pairwise preference learning is central to machine learning, with recent applications in aligning language models with human preferences. A typical dataset consists of triplets $(x…
Online Learnability of Chain-of-Thought Verifiers: Soundness and Completeness Trade-offs
Maria-Florina Balcan, Avrim Blum, Kiriaki Fragkia +2
Large Language Models (LLMs) using chain-of-thought reasoning have demonstrated great potential for solving complex reasoning and planning tasks. However, their outputs remain unre…
Learning in Structured Stackelberg Games
Maria-Florina Balcan, Kiriaki Fragkia, Keegan Harris
We initiate the study of structured Stackelberg games, a novel form of strategic interaction between a leader and a follower where contextual information can be predictive of the f…
Bicriteria Multidimensional Mechanism Design with Side Information
Maria-Florina Balcan, Siddharth Prasad, Tuomas Sandholm
We develop a versatile methodology for multidimensional mechanism design that incorporates side information about agents to generate high welfare and high revenue simultaneously. S…