6 papers
Shieldstral
Antonia Calvi, Avinash Sooriyarachchi, Giada Pistilli +274
We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7 its size on text safety benchmarks and set…
Stability and Convergence of Optimistic Exponential Weights with Asymmetric Step Sizes in Bimatrix Games
Hédi Hadiji, Sarah Sachs
We study bimatrix two-player games and investigate the last-iterate convergence and stability of equilibria for the iterates generated by the optimistic exponential weights method.…
Tracking solutions of time-varying variational inequalities
Hédi Hadiji, Sarah Sachs, Cristóbal Guzmán
Tracking the solution of time-varying variational inequalities is an important problem with applications in game theory, optimization, and machine learning. Existing work considers…
Tractable Instances of Bilinear Maximization: Implementing LinUCB on Ellipsoids
Raymond Zhang, Hédi Hadiji, Richard Combes
We consider the maximization of over , with convex and an ellipsoid. This…
Accelerated Rates between Stochastic and Adversarial Online Convex Optimization
Sarah Sachs, Hedi Hadiji, Tim van Erven +1
Stochastic and adversarial data are two widely studied settings in online learning. But many optimization tasks are neither i.i.d. nor fully adversarial, which makes it of fundamen…
Linear Bandits on Ellipsoids: Minimax Optimal Algorithms
Raymond Zhang, Hedi Hadiji, Richard Combes
We consider linear stochastic bandits where the set of actions is an ellipsoid. We provide the first known minimax optimal algorithm for this problem. We first derive a novel infor…