5 papers
Can Agentic AI Match the Performance of Human Data Scientists?
An Luo, Jin Du, Fangqiao Tian +9
Data science plays a critical role in transforming complex data into actionable insights across numerous domains. Recent developments in large language models (LLMs) have significa…
AssistedDS: Benchmarking How External Domain Knowledge Assists LLMs in Automated Data Science
An Luo, Xun Xian, Jin Du +12
Large language models (LLMs) have advanced the automation of data science workflows. Yet it remains unclear whether they can critically leverage external domain knowledge as human…
Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs
Ganghua Wang, Zhaorun Chen, Bo Li +1
As foundation models continue to scale, the size of trained models grows exponentially, presenting significant challenges for their evaluation. Current evaluation practices involve…
An Outlook on the Opportunities and Challenges of Multi-Agent AI Systems
Fangqiao Tian, An Luo, Jin Du +12
A multi-agent AI system (MAS) is composed of multiple autonomous agents that interact, exchange information, and make decisions based on internal generative models. Recent advances…
Drift to Remember
Jin Du, Xinhe Zhang, Hao Shen +7
Lifelong learning in artificial intelligence (AI) aims to mimic the biological brain's ability to continuously learn and retain knowledge, yet it faces challenges such as catastrop…