2 papers
cs.CR2026
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity
Yongxi Zhou, Junwei Yao, Yuanzhe Liu +4
A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent…
cs.SE2026
PerfAgent: Profiler-Guided Iterative Refinement for Repository-Level Code Optimization
Ryan Deng, Yuanzhe Liu, Bastian Lipka +4
Large language model (LLM) agents now perform well on correctness-oriented repository-level tasks, including SWE-Bench issue resolution and feature implementation in real codebases…