4 papers
Towards Learning Representations of Policies in Two-Player Zero-Sum Imperfect-Information Games
Kevin Wang, Kevin Yang, Arjun Prakash +1
We investigate the problem of learning useful policy representations (embeddings) in two-player zero-sum imperfect-information games. We make three contributions: First, we introdu…
Spectral Collapse Drives Loss of Plasticity in Deep Continual Learning
Arjun Prakash, Naicheng He, Kaicheng Guo +5
We investigate why deep neural networks suffer from loss of plasticity in continual learning, and thus fail to learn new tasks without reinitializing parameters. We show that this…
Distilling Game Code World Model Generation into Lightweight Large Language Models
Tyrone Serapio, Arjun Prakash, Haoyang Xu +2
Large Language Models (LLMs) have shown great ability in generating executable code from natural language, opening the possibility of automatically constructing environments for AI…
Bi-Level Policy Optimization with Nyström Hypergradients
Arjun Prakash, Naicheng He, Denizalp Goktas +2
The dependency of the actor on the critic in actor-critic (AC) reinforcement learning means that AC can be characterized as a bilevel optimization (BLO) problem, also called a Stac…