18 citations · 18 across the 2 of their papers we have counts for
3 papers
cs.AI2026
Simulating the Evolution of Alignment and Values in Machine Intelligence
Jonathan Elsworth Eicher
Model alignment is currently applied in a vacuum, evaluated primarily through standardised benchmark performance. The purpose of this study is to examine the effects of alignment o…
cs.LG2026★ 18 cited
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
cs.CL2024
Reducing Selection Bias in Large Language Models
J. E. Eicher, R. F. IrgoliÄ
Large Language Models (LLMs) like gpt-3.5-turbo-0613 and claude-instant-1.2 are vital in interpreting and executing semantic tasks. Unfortunately, these models' inherent biases adv…