3 papers
cs.SE2026
CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
Hanjun Luo, Chiming Ni, Jiaheng Wen +9
LLM-powered coding agents are reshaping the development paradigm. However, existing evaluation systems, neither traditional tests for humans nor benchmarks for LLMs, fail to captur…
cs.SD2026
AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models
Kai Li, Can Shen, Yile Liu +31
The rapid development and widespread adoption of Audio Large Language Models (ALLMs) demand rigorous evaluation of their trustworthiness. However, existing evaluation frameworks ar…
cs.CR2025
Patronus: Safeguarding Text-to-Image Models against White-Box Adversaries
Xinfeng Li, Shengyuan Pang, Jialin Wu +5
Text-to-image (T2I) models, though exhibiting remarkable creativity in image generation, can be exploited to produce unsafe images. Existing safety measures, e.g., content moderati…