activity
20162024
most citedA Gradient-based Approach for Online Robust Deep Neural Network Training with Noisy Labels

2 citations · 5 across the 6 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG20241 cited

SAIL: Self-Improving Efficient Online Alignment of Large Language Models

Mucong Ding, Souradip Chakraborty, Vibhu Agrawal +5

Reinforcement Learning from Human Feedback (RLHF) is a key method for aligning large language models (LLMs) with human preferences. However, current offline alignment approaches li…

cs.LG20232 cited

A Gradient-based Approach for Online Robust Deep Neural Network Training with Noisy Labels

Yifan Yang, Alec Koppel, Zheng Zhang

Learning with noisy labels is an important topic for scalable training in many real-world scenarios. However, few previous research considers this problem in the online setting, wh…

cs.LG20231 cited

Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General Utilities

Donghao Ying, Yunkai Zhang, Yuhao Ding +2

We investigate safe multi-agent reinforcement learning, where agents seek to collectively maximize an aggregate sum of local objectives while satisfying their own safety constraint…

cs.LG2023

Beyond Exponentially Fast Mixing in Average-Reward Reinforcement Learning via Multi-Level Monte Carlo Actor-Critic

Wesley A. Suttle, Amrit Singh Bedi, Bhrij Patel +3

Many existing reinforcement learning (RL) methods employ stochastic gradient iteration on the back end, whose stability hinges upon a hypothesis that the data-generating process mi…

cs.LG20221 cited

Online, Informative MCMC Thinning with Kernelized Stein Discrepancy

Cole Hawkins, Alec Koppel, Zheng Zhang

A fundamental challenge in Bayesian inference is efficient representation of a target distribution. Many non-parametric approaches do so by sampling a large number of points using…