3 papers
cs.CY2026
Measuring Agents in Production
Melissa Z. Pan, Negar Arabzadeh, Riccardo Cogo +22
LLM-based agents already operate in production across many industries, yet we lack an understanding of what technical methods make deployments successful. We present the first syst…
cs.CL2026
Evaluating Commercial AI Chatbots as News Intermediaries
Mirac Suzgun, Emily Shen, Federico Bianchi +5
AI chatbots are rapidly shaping how people encounter the news, yet no prior study has systematically measured how accurately these systems, with their proprietary search integratio…
cs.AI2024
Class-RAG: Real-Time Content Moderation with Retrieval Augmented Generation
Jianfa Chen, Emily Shen, Trupti Bavalatti +10
Robust content moderation classifiers are essential for the safety of Generative AI systems. In this task, differences between safe and unsafe inputs are often extremely subtle, ma…