2 papers
cs.MA2026
MasDrift: Benchmarking Authorization Preservation Across Multi-Agent Architectures
Zhuoning Xu, Xiucheng Zhang, Hanjun Luo +3
Multi-agent systems (MAS) decompose long-horizon tasks across supervisors and subagents, but delegated goals do not necessarily carry their original authorization boundaries. Exist…
cs.SD2026
AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models
Kai Li, Can Shen, Yile Liu +31
The rapid development and widespread adoption of Audio Large Language Models (ALLMs) demand rigorous evaluation of their trustworthiness. However, existing evaluation frameworks ar…