Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Single-Query Black-Box Calibration Auditing via Logit Bias
Roman Plaud, Antoine Saillenfest, Matthieu Labeau +2
Evaluating the calibration of Large Language Models (LLMs) is critical for their safe deployment as zero-shot classifiers. Yet, commercial API providers increasingly hide the conti…
cs.LG2026
Tailoring Strictly Proper Scoring Rules for Downstream Tasks: An Application to Causal Inference
Roman Plaud, Alexandre Perez-Lebel, Antoine Saillenfest +4
Probabilistic models are typically trained using task-agnostic objectives like log-loss, which can lead to significant errors in downstream estimation. This disconnect is especiall…
cs.LG2025
To Each Metric Its Decoding: Post-Hoc Optimal Decision Rules of Probabilistic Hierarchical Classifiers
Roman Plaud, Alexandre Perez-Lebel, Matthieu Labeau +2
Hierarchical classification offers an approach to incorporate the concept of mistake severity by leveraging a structured, labeled hierarchy. However, decoding in such settings freq…