7 papers
Distributionally Robust Listwise Preference Optimization
Xudong Wu, Jian Qian, Pangpang Liu +2
Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the dataset, prompt, or preference-pair level. We…
On the Convergence of Self-Improving Online LLM Alignment
Xudong Wu, Pangpang Liu, Vaneet Aggarwal +1
The Self-Improving Alignment (SAIL) algorithm addresses distribution shift by reducing a bilevel formulation of the problem to an efficient, single-level method. Empirically, SAIL…
VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models
Pang Liu, Yingjie Lao
Universal adversarial attacks on aligned multimodal large language models are increasingly reported with attack success rates in the 60-80% range, suggesting the visual modality is…
Reinforcement Learning from Human Feedback: A Statistical Perspective
Pangpang Liu, Chengchun Shi, Will Wei Sun
Reinforcement learning from human feedback (RLHF) has emerged as a central framework for aligning large language models (LLMs) with human preferences. Despite its practical success…
Fairness-aware Contextual Dynamic Pricing with Strategic Buyers
Pangpang Liu, Will Wei Sun
Contextual pricing strategies are prevalent in online retailing, where the seller adjusts prices based on products' attributes and buyers' characteristics. Although such strategies…
Uncertainty Quantification for Large Language Model Reward Learning under Heterogeneous Human Feedback
Pangpang Liu, Junwei Lu, Will Wei Sun
We study estimation and statistical inference for reward models used in aligning large language models (LLMs). A key component of LLM alignment is reinforcement learning from human…