4 papers
Toward a Principled Framework for Agent Safety Measurement
Shuyi Lin, Anshuman Suri, Alina Oprea +1
LLM agents emit actions, not just text, and once taken, those actions often cannot be undone. Yet today's agent-safety evaluations run greedy or a few sampled rollouts and report a…
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem
Shuyi Lin, Anshuman Suri, Alina Oprea +1
As large language models (LLMs) become increasingly deployed in safety-critical applications, the lack of systematic methods to assess their vulnerability to jailbreak attacks pres…
Quant Fever, Reasoning Blackholes, Schrodinger's Compliance, and More: Probing GPT-OSS-20B
Shuyi Lin, Tian Lu, Zikai Wang +3
OpenAI's GPT-OSS family provides open-weight language models with explicit chain-of-thought (CoT) reasoning and a Harmony prompt format. We summarize an extensive security evaluati…
Specification Generation for Neural Networks in Systems
Isha Chaudhary, Shuyi Lin, Cheng Tan +1
Specifications - precise mathematical representations of correct domain-specific behaviors - are crucial to guarantee the trustworthiness of computer systems. With the increasing d…