6 papers
Tracking the Best Strategy in an Extensive-Form Game
Stephen Pasteris, Rahul Savani, Theodore Turocy
We consider the extensive-form bandit problem where on each trial the learner plays an extensive-form game against an oblivious adversary. We focus on the notion of switching regre…
Differential Privacy in the Extensive-Form Bandit Problem
Stephen Pasteris, Rahul Savani, Theodore Turocy
We consider the extensive-form bandit problem, where on each trial the learner (a user coordinated by a server) plays an extensive-form game against an oblivious adversary, observi…
Guidelines for Applying RL and MARL in Cybersecurity Applications
Vasilios Mavroudis, Gregory Palmer, Sara Farmer +6
Reinforcement Learning (RL) and Multi-Agent Reinforcement Learning (MARL) have emerged as promising methodologies for addressing challenges in automated cyber defence (ACD). These…
Online Convex Optimisation: The Optimal Switching Regret for all Segmentations Simultaneously
Stephen Pasteris, Chris Hicks, Vasilios Mavroudis +1
We consider the classic problem of online convex optimisation. Whereas the notion of static regret is relevant for stationary problems, the notion of switching regret is more appro…
Fairness with Exponential Weights
Stephen Pasteris, Chris Hicks, Vasilios Mavroudis
Motivated by the need to remove discrimination in certain applications, we develop a meta-algorithm that can convert any efficient implementation of an instance of Hedge (or equiva…
Extraction Propagation
Stephen Pasteris, Chris Hicks, Vasilios Mavroudis
Running backpropagation end to end on large neural networks is fraught with difficulties like vanishing gradients and degradation. In this paper we present an alternative architect…