5 papers
OpenAI GPT-5 System Card
Aaditya Singh, Adam Fry, Adam Perelman +483
This is the system card published alongside the OpenAI GPT-5 launch, August 2025. GPT-5 is a unified system with a smart and fast model that answers most questions, a deeper reason…
Do Large Language Model Benchmarks Test Reliability?
Joshua Vendrow, Edward Vendrow, Sara Beery +1
When deploying large language models (LLMs), it is important to ensure that these models are not only capable, but also reliable. Many benchmarks have been created to track LLMs' g…
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
Ask Your Distribution Shift if Pre-Training is Right for You
Benjamin Cohen-Wang, Joshua Vendrow, Aleksander Madry
Pre-training is a widely used approach to develop models that are robust to distribution shifts. However, in practice, its effectiveness varies: fine-tuning a pre-trained model imp…
The Journey, Not the Destination: How Data Guides Diffusion Models
Kristian Georgiev, Joshua Vendrow, Hadi Salman +2
Diffusion models trained on large datasets can synthesize photo-realistic images of remarkable quality and diversity. However, attributing these images back to the training data-th…