3 papers
cs.AI2026
Structure for Reading, Prose for Writing: Asymmetric Structural Conditioning in Multi-Agent Document Authoring
Cheng Yu, Nikhil Mathew, Zhengjie Wang
Multi-agent pipelines that author formal documents must both read a requester's forms and write against them. We report a deployed tender-response system, running an open-weights m…
cs.CY2026
How Should AI Safety Benchmarks Benchmark Safety?
Cheng Yu, Severin Engelmann, Ruoxuan Cao +2
AI safety benchmarks are pivotal for safety in advanced AI systems; however, they have significant technical, epistemic, and sociotechnical shortcomings. We present a review of 210…
cs.CY2025
Safety Degradation in AI Agents
Cheng Yu, Benedikt Stroebl, Diyi Yang +1
Despite the growing integration of retrieval-enabled AI agents into society, their safety and ethical behavior remain inadequately understood. In particular, the integration of LLM…