Publications (29)
Identifying and Mitigating the Security Risks of Generative AI
Clark Barrett, Brad Boyd, Elie Burzstein +20
Higher-Order Function Networks for Learning Composable 3D Object Representations
Eric Mitchell, Selim Engin, Volkan Isler +1
FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users
Anikait Singh, Sheryl Hsu, Kyle Hsu +5
A Critical Evaluation of AI Feedback for Aligning Large Language Models
Archit Sharma, Sedrick Keh, Eric Mitchell +3
Offline Meta-Reinforcement Learning with Advantage Weighting
Eric Mitchell, Rafael Rafailov, Xue Bin Peng +2
Pixels to Plans: Learning Non-Prehensile Manipulation by Imitating a Planner
Tarik Tosun, Eric Mitchell, Ben Eisner +6
Fast Model Editing at Scale
Eric Mitchell, Charles Lin, Antoine Bosselut +2
OpenAI GPT-5 System Card
Aaditya Singh, Adam Fry, Adam Perelman +483
Memory-Based Model Editing at Scale
Eric Mitchell, Charles Lin, Antoine Bosselut +2
RLVF: Learning from Verbal Feedback without Overgeneralization
Moritz Stephan, Alexander Khazatsky, Eric Mitchell +4
Self-Destructing Models: Increasing the Costs of Harmful Dual Uses of Foundation Models
Peter Henderson, Eric Mitchell, Christopher D. Manning +2
Q-Learning for Continuous Actions with Cross-Entropy Guided Policies
Riley Simmons-Edler, Ben Eisner, Eric Mitchell +2
Higher Order Function Networks for View Planning and Multi-View Reconstruction
Selim Engin, Eric Mitchell, Daewon Lee +2
An Emulator for Fine-Tuning Large Language Models using Small Language Models
Eric Mitchell, Rafael Rafailov, Archit Sharma +2
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
Katherine Tian, Eric Mitchell, Allan Zhou +5
Test-Time Alignment via Hypothesis Reweighting
Yoonho Lee, Jonathan Williams, Henrik Marklund +4
Online Adaptation of Language Models with a Memory of Amortized Contexts
Jihoon Tack, Jaehyung Kim, Eric Mitchell +3
Meta-Learning Online Adaptation of Language Models
Nathan Hu, Eric Mitchell, Christopher D. Manning +1
Fine-tuning Language Models for Factuality
Katherine Tian, Eric Mitchell, Huaxiu Yao +2
Enhancing Self-Consistency and Performance of Pre-Trained Language Models through Natural Language Inference
Eric Mitchell, Joseph J. Noh, Siyan Li +5
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafael Rafailov, Archit Sharma, Eric Mitchell +3
Calibrating Language Models with Adaptive Temperature Scaling
Johnathan Xie, Annie S. Chen, Yoonho Lee +2
Reward Prediction Error as an Exploration Objective in Deep RL
Riley Simmons-Edler, Ben Eisner, Daniel Yang +4
On the Opportunities and Risks of Foundation Models
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli +111
Siamese Encoding and Alignment by Multiscale Learning with Self-Supervision
Eric Mitchell, Stefan Keselj, Sergiy Popovych +2
Learning Language-Conditioned Robot Behavior from Offline Data and Crowd-Sourced Annotation
Suraj Nair, Eric Mitchell, Kevin Chen +3
DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature
Eric Mitchell, Yoonho Lee, Alexander Khazatsky +2
RECKONING: Reasoning through Dynamic Knowledge Encoding
Zeming Chen, Gail Weiss, Eric Mitchell +2
OpenAI o1 System Card
OpenAI, :, Aaron Jaech +261