activity
20242026
most citedGDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

1 citations · 1 across the 4 of their papers we have counts for

collaborators

8 papers

cs.LG2026

EVMbench: Evaluating AI Agents on Smart Contract Security

Justin Wang, Andreas Bigger, Xiaohai Xu +5

Smart contracts on public blockchains now manage large amounts of value, and vulnerabilities in these systems can lead to substantial losses. As AI agents become more capable at re…

cs.CL2026

OpenAI GPT-5 System Card

Aaditya Singh, Adam Fry, Adam Perelman +483

This is the system card published alongside the OpenAI GPT-5 launch, August 2025. GPT-5 is a unified system with a smart and fast model that answers most questions, a deeper reason…

cs.LG20251 cited

GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Tejal Patwardhan, Rachel Dias, Elizabeth Proehl +16

We introduce GDPval, a benchmark evaluating AI model capabilities on real-world economically valuable tasks. GDPval covers the majority of U.S. Bureau of Labor Statistics Work Acti…

cs.LG2025

Estimating Worst-Case Frontier Risks of Open-Weight LLMs

Eric Wallace, Olivia Watkins, Miles Wang +2

In this paper, we study the worst-case frontier risks of releasing gpt-oss. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning…

cs.CL2025

gpt-oss-120b & gpt-oss-20b Model Card

OpenAI, :, Sandhini Agarwal +124

We present gpt-oss-120b and gpt-oss-20b, two open-weight reasoning models that push the frontier of accuracy and inference cost. The models use an efficient mixture-of-expert trans…

cs.CL2025

Balancing Knowledge Delivery and Emotional Comfort in Healthcare Conversational Systems

Shang-Chi Tsai, Yun-Nung Chen

With the advancement of large language models, many dialogue systems are now capable of providing reasonable and informative responses to patients' medical conditions. However, whe…