4 papers
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents
Michael Kirchhof, Gjergji Kasneci, Enkelejda Kasneci
Large-language models (LLMs) and chatbot agents are known to provide wrong outputs at times, and it was recently found that this can never be fully prevented. Hence, uncertainty qu…
Shielded Diffusion: Generating Novel and Diverse Images using Sparse Repellency
Michael Kirchhof, James Thornton, Louis Béthune +3
The adoption of text-to-image diffusion models raises concerns over reliability, drawing scrutiny under the lens of various metrics like calibration, fairness, or compute efficienc…
Benchmarking Uncertainty Disentanglement: Specialized Uncertainties for Specialized Tasks
Bálint Mucsányi, Michael Kirchhof, Seong Joon Oh
Uncertainty quantification, once a singular task, has evolved into a spectrum of tasks, including abstained prediction, out-of-distribution detection, and aleatoric uncertainty qua…