4 papers
Calibeating Made Simple
Yurong Chen, Zhiyi Huang, Michael I. Jordan +1
We study calibeating, the problem of post-processing external forecasts online to minimize cumulative losses and match an informativeness-based benchmark. Unlike prior work, which…
How Sampling Shapes LLM Alignment: From One-Shot Optima to Iterative Dynamics
Yurong Chen, Yu He, Michael I. Jordan +1
Standard methods for aligning large language models with human preferences learn from pairwise comparisons among sampled candidate responses and regularize toward a reference polic…
Mechanism Design for LLM Fine-tuning with Multiple Reward Models
Haoran Sun, Yurong Chen, Siwei Wang +3
Fine-tuning large language models (LLMs) to aggregate multiple preferences has attracted considerable research attention. With aggregation algorithms advancing, a potential economi…
Beyond Advertising: Mechanism Design for Platform-Wide Marketing Service "QuanZhanTui"
Ningyuan Li, Zhilin Zhang, Tianyan Long +8
On e-commerce platforms, sellers typically bid for impressions from ad traffic to promote their products. However, for most sellers, the majority of their sales come from organic t…