34 citations · 132 across the 41 of their papers we have counts for
74 papers
Learning to Price with Persuasion
Maria-Florina Balcan, Tejas Pagare, Karan Singh
Motivated by modern marketplaces, where the platform or the seller routinely gathers detailed user profiles, we study a novel learning theoretic model that simultaneously involves…
Efficient Online Proportional Sampling with Applications to Smoothed Online Learning
Amirmahdi Mirfakhar, Maria-Florina Balcan, Hedyeh Beyhaghi
We study the problem of efficient online proportional sampling from a high-dimensional domain under a -smoothed adversary, where the sampling distribution is induced by a dynami…
Reward Learning from Best-of- Preference Data: Targets, Tradeoffs, and Design Principles
Rattana Pukdee, Maria-Florina Balcan, Pradeep Ravikumar
Best-of- sampling is widely used to construct pairwise preference data: candidates are drawn from a base distribution, and the best is paired with a rejected response. Despi…
Verify to Amplify: Improving Reasoning via Learned Chain-of-Thought Verification
Maria-Florina Balcan, Avrim Blum, Kiriaki Fragkia +2
Large Language Models (LLMs) using chain-of-thought have demonstrated great potential for solving complex reasoning and planning tasks. Despite these advances, LLM-generated output…
The Complexity of Proper Equilibrium in Extensive-Form and Polytope Games
Brian Hu Zhang, Ioannis Anagnostides, Kiriaki Fragkia +2
The proper equilibrium, introduced by Myerson (1978), is a classic refinement of the Nash equilibrium that has been referred to as the "mother of all refinements." For normal-form…
What Does Preference Learning Recover from Pairwise Comparison Data?
Rattana Pukdee, Maria-Florina Balcan, Pradeep Ravikumar
Pairwise preference learning is central to machine learning, with recent applications in aligning language models with human preferences. A typical dataset consists of triplets $(x…