activity
20182025
most citedGeneralization in Reinforcement Learning with Selective Noise Injection and Information Bottleneck

58 citations · 152 across the 14 of their papers we have counts for

collaborators
Showing cs.LGShow all

13 papers · 1 filter

cs.LG2025

Measuring Uncertainty Calibration

Kamil Ciosek, Nicolò Felicioni, Sina Ghiassian +6

We make two contributions to the problem of estimating the calibration error of a binary classifier from a finite dataset. First, we provide an upper bound for any classifier…

cs.LG2025

Gradient Prediction with Control Variates in the Cheap-Forward Regime

Kamil Ciosek, Nicolò Felicioni, Juan Elenter +1

We study whether otherwise-idle inference resources could reduce the scarce-GPU cost of training. Our analysis uses a simulated compute ledger in which fleet work is billed at a fr…

cs.LG2025

Hallucination Detection on a Budget: Efficient Bayesian Estimation of Semantic Entropy

Kamil Ciosek, Nicolò Felicioni, Sina Ghiassian

Detecting whether an LLM hallucinates is an important research challenge. One promising way of doing so is to estimate the semantic entropy (Farquhar et al., 2024) of the distribut…

cs.LG2024

On the Importance of Uncertainty in Decision-Making with Large Language Models

Nicolò Felicioni, Lucas Maystre, Sina Ghiassian +1

We investigate the role of uncertainty in decision-making problems with natural language as input. For such tasks, using Large Language Models as agents has become the norm. Howeve…

cs.LG2023★ 14 cited

Impatient Bandits: Optimizing Recommendations for the Long-Term Without Delay

Thomas M. McDonald, Lucas Maystre, Mounia Lalmas +2

Recommender systems are a ubiquitous feature of online platforms. Increasingly, they are explicitly tasked with increasing users' long-term satisfaction. In this context, we study…

cs.LG2023

A Strong Baseline for Batch Imitation Learning

Matthew Smith, Lucas Maystre, Zhenwen Dai +1

Imitation of expert behaviour is a highly desirable and safe approach to the problem of sequential decision making. We provide an easy-to-implement, novel algorithm for imitation l…