activity
20242026
most citedOpenAI GPT-5 System Card

17 citations · 17 across the 1 of their papers we have counts for

collaborators

9 papers

cs.CL202617 cited

OpenAI GPT-5 System Card

Aaditya Singh, Adam Fry, Adam Perelman +483

This is the system card published alongside the OpenAI GPT-5 launch, August 2025. GPT-5 is a unified system with a smart and fast model that answers most questions, a deeper reason…

cs.AI2026

IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs

Chuan Guo, Juan Felipe Ceron Uribe, Sicheng Zhu +10

Instruction hierarchy (IH) defines how LLMs prioritize system, developer, user, and tool instructions under conflict, providing a concrete, trust-ordered policy for resolving instr…

cs.AI2026

FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks

Miles Wang, Robi Lin, Kat Hu +4

We introduce FrontierScience, a benchmark evaluating expert-level scientific reasoning in frontier language models. Recent model progress has nearly saturated existing science benc…

cs.AI2025

Monitoring Monitorability

Melody Y. Guan, Miles Wang, Micah Carroll +9

Observability into the decision making of modern AI systems may be required to safely deploy increasingly capable agents. Monitoring the chain-of-thought (CoT) of today's reasoning…

cs.LG2025

Persona Features Control Emergent Misalignment

Miles Wang, Tom Dupré la Tour, Olivia Watkins +8

Understanding how language models generalize behaviors from their training to a broader deployment distribution is an important problem in AI safety. Betley et al. discovered that…

cs.LG2025

Estimating Worst-Case Frontier Risks of Open-Weight LLMs

Eric Wallace, Olivia Watkins, Miles Wang +2

In this paper, we study the worst-case frontier risks of releasing gpt-oss. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning…