10 papers
Two AI Metrics Diverged: Will it Make All the Difference?
Alex Fogelson, Zachary A. Brown, Hans Gundlach +2
As exponential compute scaling continues, will the capabilities of frontier AI models outstrip what is accessible to developers on a small fixed budget? Or will capabilities conver…
EvilGenie: A Reward Hacking Benchmark
Jonathan Gabor, Jayson Lynch, Jonathan Rosenfeld
We introduce EvilGenie, a benchmark for reward hacking in programming settings. We source problems from LiveCodeBench and create an environment in which agents can easily reward ha…
The Impact of Approximation on Algorithmic Progress
Jeffery Li, Jayson Lynch, Liva Olina +3
In nearly every discipline, scientific computations are limited by the cost and speed of computation. For example, the best-known exact algorithms for the canonical Traveling Sales…
The Price of Progress: Price Performance and the Future of AI
Hans Gundlach, Jayson Lynch, Matthias Mertens +1
Language models have seen enormous progress on advanced benchmarks in recent years, but much of this progress has only been possible by using more costly models. Benchmarks may the…
How fast are algorithms reducing the demands on memory? A survey of progress in space complexity
Hayden Rome, Jayson Lynch, Jeffery Li +2
Algorithm research focuses primarily on how many operations processors need to do (time complexity). But for many problems, both the runtime and energy used are dominated by memory…
On the Origin of Algorithmic Progress in AI
Hans Gundlach, Alex Fogelson, Jayson Lynch +4
Algorithms have been estimated to increase AI training FLOP efficiency by a factor of 22,000 between 2012 and 2023 [Ho et al., 2024]. Running small-scale ablation experiments on ke…