1 paper
Aparna Elangovan, Lei Xu, Mahsa Elyasi +8
Benchmarking the capabilities of AI systems, including Large Language Models (LLMs) and Vision Models, typically ignores the impact of uncertainty in the underlying ground truth an…