activity
20162023
most citedTheoretically Principled Trade-off between Robustness and Accuracy

921 citations · 1k across the 30 of their papers we have counts for

collaborators
Showing cs.LGShow all

25 papers · 1 filter

cs.LG2023★ 4 cited

Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment

Tianhao Wu, Banghua Zhu, Ruoyu Zhang +3

Large Language Models (LLMs) can acquire extensive world knowledge through pre-training on large corpora. However, due to exposure to low-quality data, LLMs may exhibit harmful beh…

cs.LG2023★ 1 cited

On Optimal Caching and Model Multiplexing for Large Model Inference

Banghua Zhu, Ying Sheng, Lianmin Zheng +3

Large Language Models (LLMs) and other large foundation models have achieved noteworthy success, but their size exacerbates existing resource consumption and latency challenges. In…

cs.LG2023

Doubly Robust Self-Training

Banghua Zhu, Mingyu Ding, Philip Jacobson +4

Self-training is an important technique for solving semi-supervised learning problems. It leverages unlabeled data by generating pseudo-labels and combining them with a limited lab…

cs.LG2023★ 1 cited

Importance Weighted Actor-Critic for Optimal Conservative Offline Reinforcement Learning

Hanlin Zhu, Paria Rashidinejad, Jiantao Jiao

We propose A-Crab (Actor-Critic Regularized by Average Bellman error), a new practical algorithm for offline reinforcement learning (RL) in complex environments with insufficient d…

cs.LG2023★ 2 cited

Online Learning in Stackelberg Games with an Omniscient Follower

Geng Zhao, Banghua Zhu, Jiantao Jiao +1

We study the problem of online learning in a two-player decentralized cooperative Stackelberg game. In each round, the leader first takes an action, followed by the follower who ta…

cs.LG2023★ 16 cited

Principled Reinforcement Learning with Human Feedback from Pairwise or -wise Comparisons

Banghua Zhu, Jiantao Jiao, Michael I. Jordan

We provide a theoretical framework for Reinforcement Learning with Human Feedback (RLHF). Our analysis shows that when the true reward function is linear, the widely used maximum l…