Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations
Shuaiqi Wang, Aadyaa Maddi, Zinan Lin +1
Today, tool-calling agents are commonly evaluated or tested on static datasets of execution traces, including input commands, agent responses, and associated tool calls. However, i…
cs.CL2025
MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models
Chejian Xu, Jiawei Zhang, Zhaorun Chen +22
Multimodal foundation models (MMFMs) play a crucial role in various applications, including autonomous driving, healthcare, and virtual assistants. However, several studies have re…