7 papers
Fast Adversarial Attacks with Gradient Prediction
Kamil Ciosek, Aleksandr V. Petrov, Nicolò Felicioni +1
Generating adversarial examples at scale is a core primitive for robustness evaluation, adversarial training, and red-teaming, yet even "fast" attacks such as FGSM remain throughpu…
The Minimax Rate of Second-Order Calibration
Kamil Ciosek, Banafsheh Rafiee, Sina Ghiassian +1
We characterize the minimax rate of estimating the second-order calibration error for binary classification, which quantifies whether a higher-order predictor's epistemic-uncertain…
A Bayesian Information-Theoretic Approach to Data Attribution
Dharmesh Tailor, Nicolò Felicioni, Kamil Ciosek
Training Data Attribution (TDA) seeks to trace model predictions back to influential training examples, enhancing interpretability and safety. We formulate TDA as a Bayesian inform…
Measuring Uncertainty Calibration
Kamil Ciosek, Nicolò Felicioni, Sina Ghiassian +6
We make two contributions to the problem of estimating the calibration error of a binary classifier from a finite dataset. First, we provide an upper bound for any classifier…
Gradient Prediction with Control Variates in the Cheap-Forward Regime
Kamil Ciosek, Nicolò Felicioni, Juan Elenter +1
We study whether otherwise-idle inference resources could reduce the scarce-GPU cost of training. Our analysis uses a simulated compute ledger in which fleet work is billed at a fr…
Hallucination Detection on a Budget: Efficient Bayesian Estimation of Semantic Entropy
Kamil Ciosek, Nicolò Felicioni, Sina Ghiassian
Detecting whether an LLM hallucinates is an important research challenge. One promising way of doing so is to estimate the semantic entropy (Farquhar et al., 2024) of the distribut…