3 papers
cs.LG2026
Safety by Design: Realized-Cost Constraints for Contextual Bandits with Continuous Actions
Spyros Dragazis, Aldo Pacchiano
Contextual bandits are a standard framework for sequential decision-making under uncertainty, with applications in clinical trials, dosage selection, recommendation systems, and au…
stat.ML2026
A Diffusion Approximation for Temporal-Difference Learning with Linear Features under Markovian Noise
M. Forzo, E. Monzio Compagnoni, A. Russo +1
Temporal difference (TD) learning with linear function approximation is a core method for policy evaluation. Its classical continuous-time description is an ordinary differential e…
cs.LG2025
The Good, the Bad, and the Sampled: a No-Regret Approach to Safe Online Classification
Tavor Z. Baharav, Spyros Dragazis, Aldo Pacchiano
We study sequential testing for a binary disease outcome when risk follows an unknown logistic model. At each round, the decision maker may either pay for a test revealing the true…