activity
20242026
most citedAdvantage Alignment Algorithms

1 citations · 1 across the 17 of their papers we have counts for

collaborators
Showing cs.LGShow all

11 papers · 1 filter

cs.LG2026

One More Time: Revisiting Neural Quantum States from a Reinforcement Learning Perspective

Juan Agustín Duque, Sergio García Heredia, Vinicius Hernandes +4

Neural quantum states (NQS) provide a flexible and scalable framework for approximating quantum many-body wavefunctions. Among NQS parameterizations, autoregressive models are espe…

cs.LG2026

Representation Learning Enables Scalable Multitask Deep Reinforcement Learning

Johan Obando-Ceron, Lu Li, Scott Fujimoto +3

Scaling reinforcement learning (RL) to diverse multitask settings remains a central challenge. While recent advances in model-based RL achieve strong performance, they rely on plan…

cs.LG2026

Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions

Bingxu Liu, Jiashun Liu, Johan Obando-Ceron +5

While Proximal Policy Optimization (PPO) demonstrates strong performance in stationary settings, we show that its standard optimization paradigm struggles in continual and non-stat…

cs.LG2026

Towards Sustainable Investment Policies Informed by Opponent Shaping

Juan Agustin Duque, Razvan Ciuca, Ayoub Echchahed +2

Addressing climate change requires global coordination, yet rational economic actors often prioritize immediate gains over collective welfare, resulting in social dilemmas. InvestE…

cs.LG2025

Learning Robust Social Strategies with Large Language Models

Dereck Piche, Mohammed Muqeeth, Milad Aghajohari +3

As agentic AI becomes more widespread, agents with distinct and possibly conflicting goals will interact in complex ways. These multi-agent interactions pose a fundamental challeng…

cs.LG2025

Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction

Yusong Wu, Stephen Brade, Aleksandra Teng Ma +6

Most applications of generative AI involve a sequential interaction in which a person inputs a prompt and waits for a response, and where reaction time and adaptivity are not impor…