3 papers
cs.AI2026
A rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning tasks
Samiha A. Ismail, Fan X. Chen, Ali Merali
Multiple-choice medical benchmarks are increasingly saturated, and recent rubric-based evaluations such as HealthBench have shown that open-ended clinical performance is far from s…
econ.GN2025
Scaling Laws for Economic Productivity: Experimental Evidence in LLM-Assisted Consulting, Data Analyst, and Management Tasks
Ali Merali
This paper derives `Scaling Laws for Economic Impacts' -- empirical relationships between the training compute of Large Language Models (LLMs) and professional productivity. In a p…
econ.GN2024
Scaling Laws for Economic Productivity: Experimental Evidence in LLM-Assisted Translation
Ali Merali
This paper derives "scaling laws"--empirical relationships between the training compute of Large Language Models (LLMs) and their performance--for economic outcomes. In a preregist…