activity
20122023
most citedScalable trust-region method for deep reinforcement learning using Kronecker-factored approximation

471 citations · 1.1k across the 23 of their papers we have counts for

collaborators
Showing cs.LGShow all

35 papers · 1 filter

cs.LG202326 cited

Studying Large Language Model Generalization with Influence Functions

Roger Grosse, Juhan Bae, Cem Anil +14

When trying to gain better visibility into a machine learning model in order to understand and mitigate the associated risks, a potentially valuable source of evidence is: which tr…

cs.LG20223 cited

Path Independent Equilibrium Models Can Better Exploit Test-Time Computation

Cem Anil, Ashwini Pokle, Kaiqu Liang +5

Designing networks capable of attaining better performance with an increased inference budget is important to facilitate generalization to harder problem instances. Recent efforts…

cs.LG20221 cited

Proximal Learning With Opponent-Learning Awareness

Stephen Zhao, Chris Lu, Roger Baker Grosse +1

Learning With Opponent-Learning Awareness (LOLA) (Foerster et al. [2018a]) is a multi-agent reinforcement learning algorithm that typically learns reciprocity-based cooperation in…

cs.LG202248 cited

Toy Models of Superposition

Nelson Elhage, Tristan Hume, Catherine Olsson +13

Neural networks often pack many unrelated concepts into a single neuron - a puzzling phenomenon known as 'polysemanticity' which makes interpretability much more challenging. This…

cs.LG202212 cited

If Influence Functions are the Answer, Then What is the Question?

Juhan Bae, Nathan Ng, Alston Lo +2

Influence functions efficiently estimate the effect of removing a single training data point on a model's learned parameters. While influence estimates align well with leave-one-ou…

cs.LG20222 cited

Amortized Proximal Optimization

Juhan Bae, Paul Vicol, Jeff Z. HaoChen +1

We propose a framework for online meta-optimization of parameters that govern optimization, called Amortized Proximal Optimization (APO). We first interpret various existing neural…