GPT-4o System Card
arXiv:2410.21276
Abstract
GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's trained end-to-end across text, vision, and audio, meaning all inputs and outputs are processed by the same neural network. GPT-4o can respond to audio inputs in as little as 232 milliseconds, with an average of 320 milliseconds, which is similar to human response time in conversation. It matches GPT-4 Turbo performance on text in English and code, with significant improvement on text in non-English languages, while also being much faster and 50\% cheaper in the API. GPT-4o is especially better at vision and audio understanding compared to existing models. In line with our commitment to building AI safely and consistent with our voluntary commitments to the White House, we are sharing the GPT-4o System Card, which includes our Preparedness Framework evaluations. In this System Card, we provide a detailed look at GPT-4o's capabilities, limitations, and safety evaluations across multiple categories, focusing on speech-to-speech while also evaluating text and image capabilities, and measures we've implemented to ensure the model is safe and aligned. We also include third-party assessments on dangerous capabilities, as well as discussion of potential societal impacts of GPT-4o's text and vision capabilities.
Cited by in corpus (11)
- Learning from models beyond fine-tuning
- A Survey of Large Language Models for Arabic Language and its Dialects
- ImprovMate: Multimodal AI Assistant for Improv Actor Training
- GenAI Voice Mode in Programming Education
- Perovskite-R1: a domain-specialized large language model for intelligent discovery of precursor additives and experimental design
- Annotating Errors in English Learners' Written Language Production: Advancing Automated Written Feedback Systems
- Fidelity-preserving enhancement of ptychography with foundational text-to-image models
- A Roadmap for Tamed Interactions with Large Language Models
- RealHD: A High-Quality Dataset for Robust Detection of State-of-the-Art AI-Generated Images
- Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation
- Multi-agent Self-triage System with Medical Flowcharts