7 papers
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity
Yongxi Zhou, Junwei Yao, Yuanzhe Liu +4
A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent…
Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks
Yongxi Zhou, Lai Yun Choi, Jiaxi Wen +1
Run-level pass rate overstates retry-free coverage by up to 17.8 percentage points -- and the gap is largest precisely for mid-performing systems. We investigate this accuracy--sta…
Unsupervised Elicitation of Language Models
Jiaxin Wen, Zachary Ankner, Arushi Somani +10
To steer pretrained language models for downstream tasks, today's post-training paradigm relies on humans to specify desired behaviors. However, for models with superhuman capabili…
Efficient Hyperparameter Tuning via Trajectory Invariance Principle
Bingrui Li, Jiaxin Wen, Zhanpeng Zhou +2
As hyperparameter tuning becomes increasingly costly at scale, efficient tuning methods are essential. Yet principles for guiding hyperparameter tuning remain limited. In this work…
Measuring Harmfulness of Computer-Using Agents
Aaron Xuxiang Tian, Ruofan Zhang, Janet Tang +3
Computer-using agents (CUAs), which can autonomously control computers to perform multi-step actions, might pose significant safety risks if misused. However, existing benchmarks m…
Predicting Empirical AI Research Outcomes with Language Models
Jiaxin Wen, Chenglei Si, Yueh-han Chen +2
Many promising-looking ideas in AI research fail to deliver, but their validation takes substantial human labor and compute. Predicting an idea's chance of success is thus crucial…