activity
20212024
most citedScaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

20 citations · 87 across the 15 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG20243 cited

Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Jesse Farebrother, Jordi Orbay, Quan Vuong +9

Value functions are a central component of deep reinforcement learning (RL). These functions, parameterized by neural networks, are trained using a mean squared error regression ob…

cs.LG20242 cited

ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Yifei Zhou, Andrea Zanette, Jiayi Pan +2

A broad use case of large language models (LLMs) is in goal-directed decision-making tasks (or "agent" tasks), where an LLM needs to not just generate completions for a given promp…

cs.LG20232 cited

Latent Conservative Objective Models for Data-Driven Crystal Structure Prediction

Han Qi, Xinyang Geng, Stefano Rando +3

In computational chemistry, crystal structure prediction (CSP) is an optimization problem that involves discovering the lowest energy stable crystal structure for a given chemical…

cs.LG20232 cited

Beyond Uniform Sampling: Offline Reinforcement Learning with Imbalanced Datasets

Zhang-Wei Hong, Aviral Kumar, Sathwik Karnik +6

Offline policy learning is aimed at learning decision-making policies using existing datasets of trajectories without collecting additional data. The primary motivation for using r…

cs.LG20239 cited

Efficient Deep Reinforcement Learning Requires Regulating Overfitting

Qiyang Li, Aviral Kumar, Ilya Kostrikov +1

Deep reinforcement learning algorithms that learn policies by trial-and-error must learn from limited amounts of data collected by actively interacting with the environment. While…

cs.LG20216 cited

DR3: Value-Based Deep Reinforcement Learning Requires Explicit Regularization

Aviral Kumar, Rishabh Agarwal, Tengyu Ma +3

Despite overparameterization, deep networks trained via supervised learning are easy to optimize and exhibit excellent generalization. One hypothesis to explain this is that overpa…