2 papers
cs.CL2025
PIE: Performance Interval Estimation for Free-Form Generation Tasks
Chi-Yang Hsu, Alexander Braylan, Yiheng Su +2
Confidence estimation infers a probability for whether each model output is correct or not. While predicting such binary correctness is sensible for tasks with exact answers, free-…
cs.LG2023
A General Model for Aggregating Annotations Across Simple, Complex, and Multi-Object Annotation Tasks
Alexander Braylan, Madalyn Marabella, Omar Alonso +1
Human annotations are vital to supervised learning, yet annotators often disagree on the correct label, especially as annotation tasks increase in complexity. A strategy to improve…