3 papers
stat.ML2026
Reward Learning from Best-of- Preference Data: Targets, Tradeoffs, and Design Principles
Rattana Pukdee, Maria-Florina Balcan, Pradeep Ravikumar
Best-of- sampling is widely used to construct pairwise preference data: candidates are drawn from a base distribution, and the best is paired with a rejected response. Despi…
cs.LG2026
What Does Preference Learning Recover from Pairwise Comparison Data?
Rattana Pukdee, Maria-Florina Balcan, Pradeep Ravikumar
Pairwise preference learning is central to machine learning, with recent applications in aligning language models with human preferences. A typical dataset consists of triplets $(x…
cs.AI2025
LogiCity: Advancing Neuro-Symbolic AI with Abstract Urban Simulation
Bowen Li, Zhaoyu Li, Qiwei Du +10
Recent years have witnessed the rapid development of Neuro-Symbolic (NeSy) AI systems, which integrate symbolic reasoning into deep neural networks. However, most of the existing b…