1 paper · 1 filter
Aparna Elangovan, Lei Xu, Mahsa Elyasi +8
Benchmarking the capabilities of AI systems, including Large Language Models (LLMs) and Vision Models, typically ignores the impact of uncertainty in the underlying ground truth an…