23 citations · 43 across the 3 of their papers we have counts for
3 papers · 1 filter
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
RDumb: A simple approach that questions our progress in continual test-time adaptation
Ori Press, Steffen Schneider, Matthias Kümmerer +1
Test-Time Adaptation (TTA) allows to update pre-trained models to changing data distributions at deployment time. While early work tested these algorithms for individual fixed dist…
DeepGaze IIE: Calibrated prediction in and out-of-domain for state-of-the-art saliency modeling
Akis Linardos, Matthias Kümmerer, Ori Press +1
Since 2014 transfer learning has become the key driver for the improvement of spatial saliency prediction; however, with stagnant progress in the last 3-5 years. We conduct a large…