activity
20242026
collaborators

11 papers

cs.LG2026

Overcoming Valid Action Suppression in Unmasked Policy Gradient Algorithms

Renos Zabounidis, Roy Siegelmann, Mohamad Qadri +3

In reinforcement learning environments with state-dependent action validity, action masking consistently outperforms penalty-based handling of invalid actions, yet existing theory…

cs.LG2026

SCALAR: Learning and Composing Skills through LLM Guided Symbolic Planning and Deep RL Grounding

Renos Zabounidis, Yue Wu, Simon Stepputtis +4

LM-based agents excel when given high-level action APIs but struggle to ground language into low-level control. Prior work has LLMs generate skills or reward functions for RL, but…

cs.MA2026

Theory of Mind Guided Strategy Adaptation for Zero-Shot Coordination

Andrew Ni, Simon Stepputtis, Stefanos Nikolaidis +3

A central challenge in multi-agent reinforcement learning is enabling agents to adapt to previously unseen teammates in a zero-shot fashion. Prior work in zero-shot coordination of…

cs.LG2026

B3C: A Minimalist Approach to Offline Multi-Agent Reinforcement Learning

Woojun Kim, Katia Sycara

Overestimation arising from selecting unseen actions during policy evaluation is a major challenge in offline reinforcement learning (RL). A minimalist approach in the single-agent…

cs.AI2025

Adaptively Coordinating with Novel Partners via Learned Latent Strategies

Benjamin Li, Shuyang Shi, Lucia Romero +7

Adaptation is the cornerstone of effective collaboration among heterogeneous team members. In human-agent teams, artificial agents need to adapt to their human partners in real tim…

cs.MA2025

Fair Cooperation in Mixed-Motive Games via Conflict-Aware Gradient Adjustment

Woojun Kim, Katia Sycara

Multi-agent reinforcement learning in mixed-motive settings presents a fundamental challenge: agents must balance individual interests with collective goals, which are neither full…