papers
Publications (2)
cs.LG2026
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
cond-mat.soft2026
Round-Robin Test of a Light-Emitting Electrochemical Cell: Establishing a Reference Protocol for Quality Research
Anton Kirch, Kumar Saumya, Joan Rà fols-Ribé +26
Emerging technologies benefit from a jointly established reference protocol, which can lower the bar of entry for new researchers while serving as a calibration standard for establ…