17 citations · 17 across the 2 of their papers we have counts for
7 papers
Predicting LLM Safety Before Release by Simulating Deployment
Marcus Williams, Hannah Sheahan, Cameron Raymond +8
Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evidence about how often undesired model beha…
OpenAI GPT-5 System Card
Aaditya Singh, Adam Fry, Adam Perelman +483
This is the system card published alongside the OpenAI GPT-5 launch, August 2025. GPT-5 is a unified system with a smart and fast model that answers most questions, a deeper reason…
Reasoning Models Struggle to Control their Chains of Thought
Chen Yueh-Han, Robert McCarthy, Bruce W. Lee +5
Chain-of-thought (CoT) monitoring is a promising tool for detecting misbehaviors and understanding the motivations of modern reasoning models. However, if models can control what t…
Monitoring Monitorability
Melody Y. Guan, Miles Wang, Micah Carroll +9
Observability into the decision making of modern AI systems may be required to safely deploy increasingly capable agents. Monitoring the chain-of-thought (CoT) of today's reasoning…
Balancing Knowledge Delivery and Emotional Comfort in Healthcare Conversational Systems
Shang-Chi Tsai, Yun-Nung Chen
With the advancement of large language models, many dialogue systems are now capable of providing reasonable and informative responses to patients' medical conditions. However, whe…
Rule Based Rewards for Language Model Safety
Tong Mu, Alec Helyar, Johannes Heidecke +7
Reinforcement learning based fine-tuning of large language models (LLMs) on human preferences has been shown to enhance both their capabilities and safety behavior. However, in cas…