4 papers
Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges
Rui Yang, Shuang Huang, Junhua Liu +7
Safety benchmarks for large language models often assess the risk of a user query, although the outcome of question answering depends on whether the response violates a policy. Thi…
HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution
Wen Jiang, Mingmin Chu, Yimeng Tian +6
Self-evolving agents advance toward autonomy by optimizing their harness---prompts, skills, tools, and execution logic---based on environmental feedback. This paradigm, however, is…
SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems
Rui Yang, Junjie Xu, Zhengyu Liu +4
Safe agents can fail together. Multi-agent LLM systems (MAS) move information, state, decisions, and authority across principal boundaries, creating failures that local checks may…
Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts
Rui Yang, Yang Hong, Yichao Xu +3
Large Language Models (LLMs) can solve complex problems, but their misuse in high-risk domains can lead to severe consequences. Model providers therefore restrict assistance for po…