4 papers
Distributionally Robust Listwise Preference Optimization
Xudong Wu, Jian Qian, Pangpang Liu +2
Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the dataset, prompt, or preference-pair level. We…
On the Convergence of Self-Improving Online LLM Alignment
Xudong Wu, Pangpang Liu, Vaneet Aggarwal +1
The Self-Improving Alignment (SAIL) algorithm addresses distribution shift by reducing a bilevel formulation of the problem to an efficient, single-level method. Empirically, SAIL…
Rack Position Optimization in Large-Scale Heterogeneous Data Centers
Chang-Lin Chen, Jiayu Chen, Tian Lan +3
As rapidly growing AI computational demands accelerate the need for new hardware installation and maintenance, this work explores optimal data center resource management by balanci…
Global Convergence Guarantees for Federated Policy Gradient Methods with Adversaries
Swetha Ganesh, Jiayu Chen, Gugan Thoppe +1
Federated Reinforcement Learning (FRL) allows multiple agents to collaboratively build a decision making policy without sharing raw trajectories. However, if a small fraction of th…