3 papers
cs.AI2026
Improving Methodologies for Agentic Evaluations Across Domains: Leakage of Sensitive Information, Fraud and Cybersecurity Threats
Ee Wei Seah, Yongsen Zheng, Naga Nikshith +67
The rapid rise of autonomous AI systems and advancements in agent capabilities are introducing new risks due to reduced oversight of real-world interactions. Yet agent testing rema…
cs.CL2025
Learning to Route LLMs with Confidence Tokens
Yu-Neng Chuang, Prathusha Kameswara Sarma, Parikshit Gopalan +4
Large language models (LLMs) have demonstrated impressive performance on several tasks and are increasingly deployed in real-world applications. However, especially in high-stakes…
cs.LG2025
Machine Learning for Health symposium 2024 -- Findings track
Stefan Hegselmann, Helen Zhou, Elizabeth Healey +6
A collection of the accepted Findings papers that were presented at the 4th Machine Learning for Health symposium (ML4H 2024), which was held on December 15-16, 2024, in Vancouver,…