activity
20152026
most citedAgentic Large Language Models, a survey

59 citations · 118 across the 53 of their papers we have counts for

collaborators
Showing 2025Show all

15 papers · 1 filter

cs.AI2025

Mirror Mode in Fire Emblem: Beating Players at their own Game with Imitation and Reinforcement Learning

Yanna Elizabeth Smid, Peter van der Putten, Aske Plaat

Enemy strategies in turn-based games should be surprising and unpredictable. This study introduces Mirror Mode, a new game mode where the enemy AI mimics the personal strategy of a…

cs.AI2025

Guiding Skill Discovery with Foundation Models

Zhao Yang, Thomas M. Moerland, Mike Preuss +3

Learning diverse skills without hand-crafted reward functions could accelerate reinforcement learning in downstream tasks. However, existing skill discovery methods focus solely on…

cs.AI2025

A Benchmark Study of Deep Reinforcement Learning Algorithms for the Container Stowage Planning Problem

Yunqi Huang, Nishith Chennakeshava, Alexis Carras +4

Container stowage planning (CSPP) is a critical component of maritime transportation and terminal operations, directly affecting supply chain efficiency. Owing to its complexity, C…

cs.LG2025

A Unified Framework for Zero-Shot Reinforcement Learning

Jacopo Di Ventura, Jan Felix Kleuker, Aske Plaat +1

Zero-shot reinforcement learning (RL) has emerged as a setting for developing general agents, capable of solving downstream tasks without additional training or planning at test-ti…

cs.LG2025

Towards a Practical Understanding of Lagrangian Methods in Safe Reinforcement Learning

Lindsay Spoor, Álvaro Serra-Gómez, Aske Plaat +1

Safe reinforcement learning addresses constrained optimization problems where maximizing performance must be balanced against safety constraints, and Lagrangian methods are a widel…

cs.AI2025

Analysis of Bluffing by DQN and CFR in Leduc Hold'em Poker

Tarik Zaciragic, Aske Plaat, K. Joost Batenburg

In the game of poker, being unpredictable, or bluffing, is an essential skill. When humans play poker, they bluff. However, most works on computer-poker focus on performance metric…