5 papers
VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models
Pang Liu, Yingjie Lao
Universal adversarial attacks on aligned multimodal large language models are increasingly reported with attack success rates in the 60-80% range, suggesting the visual modality is…
Reinforcement Learning from Human Feedback: A Statistical Perspective
Pangpang Liu, Chengchun Shi, Will Wei Sun
Reinforcement learning from human feedback (RLHF) has emerged as a central framework for aligning large language models (LLMs) with human preferences. Despite its practical success…
Uncertainty Quantification for Large Language Model Reward Learning under Heterogeneous Human Feedback
Pangpang Liu, Junwei Lu, Will Wei Sun
We study estimation and statistical inference for reward models used in aligning large language models (LLMs). A key component of LLM alignment is reinforcement learning from human…
Fairness-aware Contextual Dynamic Pricing with Strategic Buyers
Pangpang Liu, Will Wei Sun
Contextual pricing strategies are prevalent in online retailing, where the seller adjusts prices based on products' attributes and buyers' characteristics. Although such strategies…
Dual Active Learning for Reinforcement Learning from Human Feedback
Pangpang Liu, Chengchun Shi, Will Wei Sun
Aligning large language models (LLMs) with human preferences is critical to recent advances in generative artificial intelligence. Reinforcement learning from human feedback (RLHF)…