works on

From the 1 of 14 linked papers with an AI index.

activity
20242026
most citedJaxMARL: Multi-Agent RL Environments and Algorithms in JAX

2 citations · 2 across the 7 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG20262 cited

JaxMARL: Multi-Agent RL Environments and Algorithms in JAX

Alexander Rutherford, Benjamin Ellis, Matteo Gallici +18

Benchmarks are crucial in the development of machine learning algorithms, with available environments significantly influencing reinforcement learning (RL) research. Traditionally,…

cs.LG2026

Tackling GNARLy Problems: Graph Neural Algorithmic Reasoning Reimagined through Reinforcement Learning

Alex Schutz, Victor-Alexandru Darvariu, Efimia Panagiotaki +2

Neural algorithmic reasoning (NAR) is a paradigm that trains neural networks to execute classic algorithms by supervised learning. Despite its successes, important limitations rema…

cs.LG2026

Improving Regret Approximation for Unsupervised Dynamic Environment Generation

Harry Mead, Bruno Lacerda, Jakob Foerster +1

Unsupervised Environment Design (UED) seeks to automatically generate training curricula for reinforcement learning (RL) agents, with the goal of improving generalisation and zero-…

cs.LG2025

JaxWildfire: A GPU-Accelerated Wildfire Simulator for Reinforcement Learning

Ufuk Çakır, Victor-Alexandru Darvariu, Bruno Lacerda +1

Artificial intelligence methods are increasingly being explored for managing wildfires and other natural hazards. In particular, reinforcement learning (RL) is a promising path tow…

cs.LG2025

Return Capping: Sample-Efficient CVaR Policy Gradient Optimisation

Harry Mead, Clarissa Costen, Bruno Lacerda +1

When optimising for conditional value at risk (CVaR) using policy gradients (PG), current methods rely on discarding a large proportion of trajectories, resulting in poor sample ef…

cs.LG2025

DITTO: Offline Imitation Learning with World Models

Branton DeMoss, Paul Duckworth, Jakob Foerster +2

For imitation learning algorithms to scale to real-world challenges, they must handle high-dimensional observations, offline learning, and policy-induced covariate-shift. We propos…