4 papers
DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle
Fangyu Lei, Jinxiang Meng, Yiming Huang +14
Real-world enterprise data intelligence workflows encompass data engineering that turns raw sources into analytical-ready tables and data analysis that convert those tables into de…
Towards Generalizable Context-aware Anomaly Detection: A Large-scale Benchmark in Cloud Environments
Xinkai Zou, Xuan Jiang, Ruikai Huang +8
Anomaly detection in cloud environments remains both critical and challenging. Existing context-level benchmarks typically focus on either metrics or logs and often lack reliable a…
SAFEFLOW: A Principled Protocol for Trustworthy and Transactional Autonomous Agent Systems
Peiran Li, Xinkai Zou, Zhuohang Wu +9
Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled powerful autonomous agents capable of complex reasoning and multi-modal tool use. Des…
Goal2Story: A Multi-Agent Fleet based on Privately Enabled sLLMs for Impacting Mapping on Requirements Elicitation
Xinkai Zou, Yan Liu, Xiongbo Shi +1
As requirements drift with rapid iterations, agile development becomes the dominant paradigm. Goal-driven Requirements Elicitation (RE) is a pivotal yet challenging task in agile p…