activity
20242026
collaborators

6 papers

cs.LG2026

Tracking the Best Strategy in an Extensive-Form Game

Stephen Pasteris, Rahul Savani, Theodore Turocy

We consider the extensive-form bandit problem where on each trial the learner plays an extensive-form game against an oblivious adversary. We focus on the notion of switching regre…

cs.CR2026

Differential Privacy in the Extensive-Form Bandit Problem

Stephen Pasteris, Rahul Savani, Theodore Turocy

We consider the extensive-form bandit problem, where on each trial the learner (a user coordinated by a server) plays an extensive-form game against an oblivious adversary, observi…

cs.AI2025

Guidelines for Applying RL and MARL in Cybersecurity Applications

Vasilios Mavroudis, Gregory Palmer, Sara Farmer +6

Reinforcement Learning (RL) and Multi-Agent Reinforcement Learning (MARL) have emerged as promising methodologies for addressing challenges in automated cyber defence (ACD). These…

cs.LG2025

Online Convex Optimisation: The Optimal Switching Regret for all Segmentations Simultaneously

Stephen Pasteris, Chris Hicks, Vasilios Mavroudis +1

We consider the classic problem of online convex optimisation. Whereas the notion of static regret is relevant for stationary problems, the notion of switching regret is more appro…

cs.LG2025

Fairness with Exponential Weights

Stephen Pasteris, Chris Hicks, Vasilios Mavroudis

Motivated by the need to remove discrimination in certain applications, we develop a meta-algorithm that can convert any efficient implementation of an instance of Hedge (or equiva…

cs.LG2024

Extraction Propagation

Stephen Pasteris, Chris Hicks, Vasilios Mavroudis

Running backpropagation end to end on large neural networks is fraught with difficulties like vanishing gradients and degradation. In this paper we present an alternative architect…