128 citations · 133 across the 3 of their papers we have counts for
3 papers
Meta-Designing Quantum Experiments with Language Models
Sören Arlt, Haonan Duan, Felix Li +3
Artificial Intelligence (AI) can solve complex scientific problems beyond human capabilities, but the resulting solutions offer little insight into the underlying physical principl…
When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Norah Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay +9
Large Language Model (LLM) leaderboards based on benchmark rankings are regularly used to guide practitioners in model selection. Often, the published leaderboard rankings are take…
Holistic Evaluation of Language Models
Percy Liang, Rishi Bommasani, Tony Lee +47
Language models (LMs) are becoming the foundation for almost all major language technologies, but their capabilities, limitations, and risks are not well understood. We present Hol…