activity
20112023
most citedJudging LLM-as-a-Judge with MT-Bench and Chatbot Arena

492 citations · 1.5k across the 62 of their papers we have counts for

collaborators
Showing 2021 · cs.LGShow all

10 papers · 2 filters

cs.LG2021

The Effect of Model Size on Worst-Group Generalization

Alan Pham, Eunice Chan, Vikranth Srivatsa +6

Overparameterization is shown to result in poor test accuracy on rare subgroups under a variety of settings where subgroup information is known. To gain a more complete picture, we…

cs.LG2021★ 4 cited

C-Planning: An Automatic Curriculum for Learning Goal-Reaching Tasks

Tianjun Zhang, Benjamin Eysenbach, Ruslan Salakhutdinov +2

Goal-conditioned reinforcement learning (RL) can solve tasks in a wide range of domains, including navigation and manipulation, but learning to reach distant goals remains a centra…

cs.LG2021★ 20 cited

Accelerating Quadratic Optimization with Reinforcement Learning

Jeffrey Ichnowski, Paras Jain, Bartolomeo Stellato +6

First-order methods for quadratic optimization such as OSQP are widely used for large-scale machine learning and embedded optimal control, where many related problems must be rapid…

cs.LG2021

Taxonomizing local versus global structure in neural network loss landscapes

Yaoqing Yang, Liam Hodgkinson, Ryan Theisen +4

Viewing neural network models in terms of their loss landscapes has a long history in the statistical mechanics approach to learning, and in recent years it has received attention…

cs.LG2021★ 2 cited

LS3: Latent Space Safe Sets for Long-Horizon Visuomotor Control of Sparse Reward Iterative Tasks

Albert Wilcox, Ashwin Balakrishna, Brijen Thananjeyan +2

Reinforcement learning (RL) has shown impressive success in exploring high-dimensional environments to learn complex tasks, but can often exhibit unsafe behaviors and require exten…

cs.LG2021★ 11 cited

MADE: Exploration via Maximizing Deviation from Explored Regions

Tianjun Zhang, Paria Rashidinejad, Jiantao Jiao +3

In online reinforcement learning (RL), efficient exploration remains particularly challenging in high-dimensional environments with sparse rewards. In low-dimensional environments,…