14 citations · 14 across the 1 of their papers we have counts for
3 papers
cs.AI2026★ 14 cited
Measuring AI Ability to Complete Long Software Tasks
Thomas Kwa, Ben West, Joel Becker +23
Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities,…
hep-ph2025
Unifying Simulation and Inference with Normalizing Flows
Haoxing Du, Claudius Krause, Vinicius Mikuni +3
There have been many applications of deep neural networks to detector calibrations and a growing number of studies that propose deep generative models as automated fast detector si…
cs.LG2025
WeatherMesh-3: Fast and accurate operational global weather forecasting
Haoxing Du, Lyna Kim, Joan Creus-Costa +5
We present WeatherMesh-3 (WM-3), an operational transformer-based global weather forecasting system that improves the state of the art in both accuracy and computational efficiency…