paper

Online Learning in Stackelberg Security Games with Adaptive Attacker Sequences and Time-Varying Attack Intensities

arXiv:2608.01703

Abstract

This work studies no-regret online learning in Repeated Stackelberg Security Games with time-varying attack intensities. We formulate an extended security game in which an attacker may select multiple targets and derive an exact mixed-integer linear programming oracle under a optimistic tie-breaking rule. Under full-information feedback, the oracle is integrated with Follow-the-Perturbed-Leader and yields expected regret against non-anticipating sequences with time-varying follower numbers, attack intensities, and attacker types. Under bandit feedback, we consider multiple followers sharing a fixed attacker type and use a barycentric-spanner construction to reconstruct utility estimates from aggregate attack observations, obtaining expected regret. Extensive simulations demonstrate the robustness and effectiveness of our approach under full and partial information feedback.

Online Learning in Stackelberg Security Games with Adaptive Attacker Sequences and Time-Varying Attack Intensities · wovepaper