3 papers
cs.AI2026
MirrorCode: AI can rebuild entire programs from behavior alone
Tom Adamczewski, David Owen, David Rein +4
AI models are rapidly improving at autonomous coding, as shown by benchmark progress and one-off demonstrations such as AI implementing a C compiler. However, existing coding bench…
cs.CY2026
The ATOM Report: Measuring the Open Language Model Ecosystem
Nathan Lambert, Florian Brand
We present a comprehensive adoption snapshot of the leading open language models and who is building them, focusing on the ~1.5K mainline open models from the likes of Alibaba's Qw…
cs.CL2025
ReadBench: Measuring the Dense Text Visual Reading Ability of Vision-Language Models
Benjamin Clavié, Florian Brand
Recent advancements in Large Vision-Language Models (VLMs), have greatly enhanced their capability to jointly process text and images. However, despite extensive benchmarks evaluat…