most citedTrading off rewards and errors in multi-armed bandits

11 citations

7 papers

cs.SC2026

Delayed Constraints in Narrowing for the Logic-Based Analyses of Real-Time Systems

Santiago Escobar, Raúl López-Rueda, Carlos Olarte

The formal analysis of real-time systems must address two dimensions of infiniteness: an unbounded number of agents and messages, and a potentially infinite state space induced by…

cs.LO2026

The Bright Side of Timed Opacity

Étienne André, Sarah Dépernet, Engel Lefaucheux

Timed automata (TAs) are an extension of finite automata that can measure and react to the passage of time, providing the ability to handle real-time constraints using clocks. In 2…

cs.IR2026

A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval Models

Yash Kankanampati, Yuxuan Zong, Nadi Tomeh +2

Late-interaction models such as ColBERT offer competitive performance across various retrieval tasks but require storing a dense embedding for each document token, leading to a sub…

cs.LO2026

Efficient Decision Procedures for RNmatrix Semantics

Renato R. Leme, Carlos Olarte, Elaine Pimentel

Logical matrices provide a semantic framework in which connectives are interpreted by deterministic truth-functions. While elegant, this approach is often too restrictive to captur…

cs.LO2026

Collusion Relations and their Applications to Balance Theory

Jean-Baptiste Joinet, Carlos Olarte

We study quadrangular properties of binary relations on a set X -i.e., properties defined on configurations of four elements -- within an agonistic interpretation, where xRy is int…

cs.LG202611 cited

Trading off rewards and errors in multi-armed bandits

Akram Erraqabi, Alessandro Lazaric, Michal Valko +2

In multi-armed bandits, the most-explored arms are the most informative, while reward maximization typically pulls only the best arm. We study the tradeoff between identifying arm…